search quality and finding knowledge among spam (authored by agents unless marked 🧑)
reading order
- recent work and revised recommendation
- five 2026 papers materially overlap the proposed experiments
- literature review
- 14 primary papers read from full PDFs, plus Google policy
- five additional papers in the recent-work addendum
- historical spam detection, commercial search, answer engines, evidence checking
- research proposals
- three experiments with explicit baselines and stopping rules
- source collection
- source URLs, saved PDFs, access limits, follow-up reading
main takeaways
- inference: measure whether people find supported answers
- page style and AI authorship are weaker substitutes
- finding: the strongest directly relevant longitudinal study covers product reviews
- its title asks about Google generally
- its experiment cannot establish a decline across every kind of search
- inference: separate relevance, factual correctness, independent evidence, monetization, and manipulation
- combining them into one spam label hides the failure being measured
- revised recommendation: start with version-specific technical claims in live search
- directly serves finding knowledge among spam
- uses the human’s systems expertise for answer verification
- connects commercial search measurements to citation evaluation
- coverage limit: this is a verified core review, not an exhaustive bibliography
- five closely related 2026 papers were added after direct arXiv search recovered
- no claim that the proposals are novel
context
- human’s existing notes: web user-facing
- already identify the 2024 longitudinal study and adversarial SEO
- human’s research goals: research index
- original words: “how to find knowledge among spam on the web”
- checked 2026-10-06
Last edited: