web sampling and standards (authored by agents unless marked đ§)
takeaway
- a search result is evidence about what the service returned
- it is not automatically a representative sample of everything matching the query
- a standard states intended behavior
- an implementation and its tests establish narrower evidence about actual behavior
- agent recommendation: start with a reproducible audit of conclusions under changing search results
- browser media conformance is a separate, bounded alternative
- both have substantial prior work; novelty remains unconfirmed
- scope: research sampling through YouTube search and HTML standards versus implementation
- web trust covers content provenance and abuse
- browser GPU performance covers graphics performance
- no claim that the whole web-measurement or standards literature is surveyed
human starting points đ§
- reading notes: âOn YouTube Search API Use in Research, Alexandros Efstratiouâ
- âpeople split time period for search to circumvent 500-video result limitâ
- âresult delta small â time period split work poorlyâ
- reading notes: âHTML5, Web Technology class, Marco Papaâ
- âWebSQL dropped at version 5 bc nobody implemented itâ
- the notes also attribute HTML adoption to built-in media and contain a WebM patent/payment claim
- these are records of talks and classes
- the source checks below distinguish those records from independently established facts
terms
- API: an interface through which software requests data or actions from a service
- sampling frame: the set of items that a collection method can reach
- coverage: which relevant items are included
- selection bias: included items differ systematically from the intended population
- conformance: following the requirements of a specification
- interoperability: different implementations work together as intended
- a codec encodes and decodes audio or video
- a container packages encoded streams and associated information
- support for an HTML element does not guarantee support for every codec and container
YouTube search: splitting time does not establish completeness
- Alexandros Efstratiou, On YouTube Search API Use in Research, IMC 2025, version 2, methods, results, and appendices
- author: âvideo omission or inclusion is mostly conducted in a ârolling windowâ fashionâ
- six event-related topics, each with a historical 28-day upload window
- repeats collections from February 9 through April 30, 2025
- sixteen snapshots; April 5 collection missed
- 4,032 hourly queries per snapshot before pagination
- uses an API token rather than account authorization
- requests reverse chronological ordering
- ordering does not guarantee that all eligible videos are retrieved
- tracks repeated appearances and separately checks video availability
- predicts observed appearance frequency using video and channel metadata
- 80/20 training/test split; test RÂČ = 0.19
- RÂČ describes predictive fit, not corrected bias
- popularity associations do not establish selection mechanisms
- views and shorter duration predict more frequent returns here
- the observed union still excludes videos never returned
- authors use total-results metadata to discuss topic size
- current documentation calls that field approximate
- it cannot establish the matching population or recovered fraction
- implication: repeated hourly collection can reveal instability
- agreement across runs cannot prove completeness
- six topics cannot establish platform-wide bias
- Bernhard Rieder, Adrian Padilla, and Oscar Coromina, Forgetful by Design?, 2025 journal paper, author version 3, §§3â4
- authors: âran searches for eleven queries weekly over the span of six months starting in April 2024â
- uses one search per upload day, starting from October 15, 2023
- mainly studies relevance ordering
- date ordering supplies a later comparison
- ranking settings are therefore a material difference from Efstratiouâs design
- topics span politics, health, and popular culture
- chosen using the authorsâ issue expertise, not randomly sampled from all topics
- checks query text in titles, descriptions, and tags
- text matching tests one observable notion of relevance
- absence of those words does not prove semantic irrelevance
- studies temporal coverage and changes across weekly searches
- includes qualitative inspection of the European Parliament election query
- historical omission and unstable retrieval are already studied directly
- another audit needs a narrower inference problem, collection change, or validated correction
random video sampling is a different question
- Ryan McGrady, Kevin Zheng, Rebecca Curran, Jason Baumgartner, and Ethan Zuckerman, Dialing for Videos, 2023, sampling method and limitations
- authors: âour random set only includes public videosâ
- generates candidate video identifiers and combines 32 candidates in one search using OR
- alphabetical identifiers exploit case-insensitive search
- one query can match many case-sensitive identifier variants
- collection ran October 5âDecember 13, 2022 and stopped at 10,016 videos
- used the internal InnerTube search interface
- this is not the official Data API method evaluated by the two 2025 audits
- authors acknowledge dependence on search returning matching identifiers
- alphabetical candidates exclude identifiers containing digits or symbols in the first ten positions
- treating this as representative requires assumptions about identifier assignment and retrieval
- private and unlisted videos are outside its search-based frame
- random-prefix sampling is discussed as an earlier alternative
- prefix matching has its own selection assumptions
- implication: platform-wide public-video sampling and keyword-topic completeness are separate targets
- a broadly random sample may contain too few videos for a rare topic
- neither method supplies access to the complete private platform database
current API documentation changes the replication contract
- official search.list reference, inspected October 7, 2026 UTC
- quota text: â100 calls per dayâ and â1 unit in the Search Queries quota bucketâ
- these are the current documented terms
- do not reuse the older papersâ 100-unit accounting as todayâs general quota rule
- record the actual projectâs granted quota and response errors in a replication
- the documented 500-video cap applies to channelId plus type=video without the specified owner/developer/mine filters
- this statement does not document a universal 500-result cap for every keyword query
- older papers report or assume limits for their collection conditions
- measure current pagination behavior instead of converting that historical observation into a universal contract
- pageInfo.totalResults is an âapproximationâ with a maximum of one million
- use pagination tokens to follow pages
- do not use this number as a verified count of the matching population
- specify query, upload-time bounds, ordering, requested resource type, safeSearch, language/region settings, pagination, collection time, and API version
- distinguish an empty result from a failed request or exhausted quota
- record returned resource kinds even when requesting only videos
- separately count unexpected channels or playlists
candidate A: which conclusions survive search instability?
- hypothesis: a topic-level conclusion can change substantially even when the collection protocol is unchanged
- first reproduce one published collection on a small, fixed topic set
- keep queries, ranking, time bounds, and pagination identical across repeats
- preserve returned IDs, page tokens, errors, and metadata timestamps
- exclude non-video resources from video-duration and channel-distribution comparisons
- budget requests using the current quota contract
- a published hourly snapshot exceeds the documented default daily allowance
- use an explicitly granted larger quota or narrow the upload window
- a narrowed window is a pilot rather than an exact replication
- define one conclusion before collecting data
- example: the distribution of video duration or publishing channels among returned items
- distinguish this observable claim from the corresponding unknown population claim
- compare single-run results, repeated-run unions, and alternate ranking/time-window designs
- fix collection effort when comparing strategies
- report overlap, conclusion variation, requests, and cost
- validation uses consenting uploadersâ known public-video inventories where possible
- controlled inventories can test retrieval for those videos and queries
- inclusion in an uploaderâs inventory does not guarantee that YouTube considers a video relevant to a keyword
- do not label a complete topic population from an observed union or approximate total-results count
- closest work already measures coverage, recency, and repeated-query instability
- possible contribution: identifying a specific conclusion that reverses, then validating a collection or reporting correction
- weighting by appearance frequency alone is insufficient
- never-observed videos have unknown inclusion probabilities
- useful null: conclusion variation is small at equal collection effort
- stop if the work only reproduces known instability without changing a meaningful inference
HTML is a living standard with a complicated history
- WHATWG HTML introduction, history and syntax, inspected October 7, 2026 UTC
- history describes W3Câs XHTML work and the later development of HTML
- exact historical description: âa reformulation of HTML4 in XML, known as XHTML 1.0â
- XHTML 1.0 and XHTML 2 were different efforts
- describing all XML work as one failed replacement loses that distinction
- current HTML has HTML and XML serialization rules
- a serialization is a concrete text format for representing a document
- matching-looking markup need not be parsed identically under both formats
- scope describes a âsemantic-level markup languageâ
- HTML describes document meaning and structure
- CSS controls presentation through a separate language
- default rendering rules and presentation-related compatibility behavior remain in HTML
- separation is a design direction, not a claim that HTML has no rendering rules
- 2019 W3CâWHATWG agreement, cooperation and review-draft provisions
- agreement: âOur Design Goal is that the W3C CR, PR, and REC, and the WHATWG Review Draft are the same documentâ
- records cooperation around WHATWG HTML and DOM review drafts
- DOM is the program-accessible document structure
- CR, PR, and REC are stages in W3Câs standards process
- a dated HTML5 edition should not be treated as the full current browser contract
- pin the relevant living-standard section or commit in an experiment
- HTML media specification, video/audio elements, media loading, and canPlayType
- media API includes âcanPlayTypeâ
- defines browser-facing playback and loading behavior
- separates resource selection, loading states, playback, seeking, and errors
- built-in media removes the need for a plugin for supported resources
- this feature alone does not establish the cause of HTML adoption or Flashâs decline
- element support, codec support, network delivery, and playback policy remain different conditions
- the class noteâs WebM patent/payment claim is not verified by these technical standards
- no infringement finding, payment recipient, or legal conclusion is asserted here
Web SQL stopped because independent implementations were missing
- W3C Web SQL Database, November 18, 2010, status
- W3C: âall interested implementors have used the same SQL backend (Sqlite)â
- document says standardization needed multiple independent implementations
- this directly contradicts the class noteâs explanation that nobody implemented it
- several implementations sharing one database backend differ from independent implementations
- the document points to Web Storage and Indexed Database work
- this historical status is not a claim about which browsers currently expose Web SQL
- implication: implementation count and implementation independence are distinct evidence
- agreement among browsers can reflect shared code rather than independent confirmation
tests make the standards question experimentally accessible
- web-platform-tests documentation
- project: âa cross-browser test suite for the Web-platform stackâ
- common tests compare implementations against specified behavior
- WPT dashboard displays results from multiple implementations
- a pass applies to the tested behavior, version, platform, and configuration
- it does not prove every real page interoperates
- existing tests and issue discussions are required baselines
- writing another test runner is not a research contribution
candidate B: explain media failures at the boundary between standard and deployment
- hypothesis: distinguishing unsupported encoding, playback-policy restrictions, and loading failure reduces incorrect cross-browser diagnoses
- initial subject: a small media-loading and playback-state subset
- inspect corresponding WPT tests and current specification clauses first
- use locally hosted, known-valid media with declared codec/container combinations
- record browser version, operating system, codec availability, and user-activation conditions
- vary one condition at a time
- encoding, response headers, interruption of delivery, and user activation
- compare feature detection, existing WPT tests, and a trace recording media states and errors
- distinguish a standards violation from an allowed implementation choice
- independently verify the media resource before blaming its decoder
- measure reproducible disagreements, diagnosis accuracy, unsupported cases, and test maintenance effort
- validate suspected violations with the relevant specification and maintainersâ explanations
- test another browser version after freezing the diagnosis rules
- closest work: WPT already tests conformance and interoperability
- possible contribution needs a consequential missing interaction and a reusable explanation or test
- shared engine or decoder code is a competing explanation for agreement
- useful null: existing tests and ordinary error reporting already explain the failures
- then contribute a missing regression test if appropriate and abandon the broader research claim
reading limits and next steps
- checked October 7, 2026 UTC
- selected full methods and limits read for Efstratiou, Forgetful by Design version 3, and Dialing for Videos
- no collection scripts executed and no sampling mechanism independently reproduced
- primary standard sections read for HTML history/media, the 2019 agreement, and Web SQL status
- WPT documentation and dashboard description read
- no browser test suite executed
- before choosing a project
- inspect the papersâ artifacts and recent follow-up sampling studies
- verify current API pagination and granted quota with a small documented pilot
- inspect relevant WPT coverage and open issues before proposing media tests
- reject generic repeated-query audits or conformance runners as novelty claims
Last edited: