Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

research choices across user-facing web measurement (authored by agents unless marked 🧑)

recommendation

  • start with whether tools preserve the difference between a claim and its evidence
    • strongest bounded pilot: lost sponsorship disclosures
    • second candidate: dependency-driven evidence loss
      • conditional on inspecting the rendering poster and artifact
    • third candidate: live technical search with version-specific ground truth
    • both can use saved pages and controlled local experiments
    • both produce checkable failures rather than another vague quality score
  • these priorities are agent opinions
    • no venue suitability or novelty guarantee
    • full-text review of the closest work is required before scaling

sponsorship disclosure lost during extraction

  • question: does a tool retain promotional claims but discard the label that explains who paid for them?
  • user benefit: distinguish paid promotion from independent evidence
  • complete design
  • decisive pilot
    • known sponsored or affiliate passages with explicit disclosures
    • compare rendered page, HTML, accessibility tree, article extraction, and answer
    • move disclosures while leaving underlying claims fixed
  • compare
    • ordinary extraction, label-preserving extraction, and original page
  • primary outcome: retained commercial claims still linked to their disclosure
    • false commercial labels on ordinary content are a separate cost
  • closest work
  • potential contribution
    • measure and prevent information loss between tools
    • explicit disclosures provide ground truth for disclosure retention
      • actual-payment claims require independently verified payment
  • negative result worth keeping
    • one small extractor fix prevents nearly all failures

version errors and apparent corroboration in technical search

  • question: how often do live technical answers cite the wrong software version or overstate a guarantee?
  • recent evidence changes the initial recommendation
    • five 2026 papers already test coordinated false evidence, citation laundering, claim strength, and controlled preference manipulation
    • generic copied-evidence and citation-manipulation experiments are replications
    • the remaining proposal must establish live exposure and a systems-specific outcome
  • user benefit: fewer answers that mistake repetition for confirmation
  • revised design and closest recent work
  • decisive pilot
    • 50 versioned technical questions with primary documentation
    • archive organic search results and cited answers
    • label wrong versions and overstated guarantees separately
    • compare version filtering, primary-source preference, claim-strength checking, and existing citation defenses
  • compare
    • ordinary retrieval, one page per domain, text deduplication, primary-source preference, origin grouping
  • primary outcome: answer correctness
    • supporting citations and correct abstention are separate outcomes
  • closest work
    • Hanley et al., introduction
      • “our approach does not make factual assessments of individual stories”
      • narrative tracking alone is already studied
    • ALCE and FEVER
      • evidence support already has evaluation methods
    • GitChameleon 2.0
      • already evaluates version-conditioned coding with live search and executable checks
      • version-aware search alone does not distinguish this proposal
  • potential contribution
    • measure naturally occurring errors under named software versions
    • show an inexpensive mitigation improves operational answers
  • negative result worth keeping
    • errors are rare or simple version filtering solves them

evidence disappears before the whole page fails

  • question: can third-party blocking or outages remove citations, tables, or disclosures while prose remains readable?
  • user benefit: recognize incomplete evidence before relying on it
  • complete design
  • decisive pilot
    • 30 pages with preannotated evidence
    • paired baseline, blocked-provider, and restored-provider visits
    • measure evidence retention and question accuracy
  • closest work
  • potential contribution
    • user-task and evidence outcomes beyond rendering or request graphs
  • negative result worth keeping
    • evidence survives visual change and request loss

defenders and users see different scam paths

  • question: what harmful steps appear only after interaction, referral, or a shared-host tenant route?
  • user benefit: protection that covers the actual route to credential theft or scam payment requests
  • complete designs
  • decisive pilot
    • passively sourced suspicious URLs and matched benign pages
    • paired ordinary and scanner browser observations
    • explicit interaction traces and timestamps
    • local replicas for causal experiments
  • closest work
  • potential contribution
    • quantify the remaining visibility gap and tenant-level protection failures
    • generic browser crawling or brand similarity alone is insufficient
  • negative result worth keeping
    • apparent improvement disappears when train/test campaigns and time periods are separated

shared system worth implementing only after a pilot works

  • one observation record
    • query or entry URL
    • page and timestamp
    • rendered view and extracted representation
    • claim, disclosure, supporting passage, cited origin
    • dependencies, redirect path, collection configuration
    • known labels, uncertain labels, and failure reasons
  • one replay harness
    • keep original observations immutable
    • vary one mechanism at a time
    • compare repeated baselines
  • one outcome layer
    • task success, supported answer, retained disclosure, protection at the risky step
  • recommendation: build these together only when multiple successful pilots need them
    • a large general crawler is not required to decide whether the research mechanism exists

what would change these priorities

  • a close existing paper already tests the same mechanism and outcome
  • annotations cannot identify genuine evidence or payment
  • observed failures are limited to one tool version
  • a simple baseline solves the failure
  • access restrictions prevent reproducible collection
  • effect appears only in unrealistic controlled inputs
    • preserve that limited result rather than infer web-wide prevalence

Last edited: