reproducing reported mobile bugs (authored by agents unless marked đ§)
takeaway
- reproduce the reported failure, not merely a plausible sequence of taps
- search finds an action sequence
- replay executes it under recorded conditions
- a failure check determines whether the reported problem occurred
- papers call this check an oracle
- broad ideas already have substantial prior work
- global search, LLM planning, screenshots, missing-step inference, and non-crash failure recognition
- promising narrower question: distinguish a reported visible failure from ordinary behavior after different UI actions
- an empty search result can be correct
- a crash can come from the wrong interaction
- novelty remains unverified
human starting point
- đ§ âDefense: Automated Reproduction of Bug Reports for Mobile Applications, Zhaoxu Zhangâ
- source: reading notes
- đ§ âignore hidden state; only consider UI stateâ
- source: same defense notes
- reading-note description of the defense, not a restriction on proposed research
- đ§ âintermediate representation to help validate non-crash bug reproductionâ
- source: same defense notes
- Zhangâs USC defense announcement, May 9 2025
- author: âautomatically recognizes buggy behaviors based on bug reportsâ
- identifies non-crash reproduction as part of the dissertation
- this announcement motivates reading the papers; it does not independently validate their results
review scope
- selected full-method and evaluation sections read
- RepRev 2020, AndroR2 2021, Android reproduction study 2022, ReproBot 2023, AdbGPT ICSE 2024, Roam 2024, ReBL 2024, image study 2025, AndroB2O ASE 2025, TreeMind February 2026 version
- implementation and prompt inspection, not full evaluation
- BugSpot FSE 2025 artifact at commit
52ab3fa96793421fc43387c81830d3e6113620de
- BugSpot FSE 2025 artifact at commit
- abstract only
- ReActDroid TSE 2025, ICST 2026 intent study, ReproPilot ASE 2026
- access checked October 6 2026
- BugSpot publisher download returned HTML rather than the paper
- no matching full paper was obtained for ReproPilot or the ICST intent study
- repository lead only
- CARBON
- no new tool was implemented or experimentally evaluated
- proposals are agent hypotheses
datasets explain what success rates mean
- Wendland et al, MSR 2021, AndroR2 paper, collection and reproduction phases
- authors: â90 manually reproduced bug reportsâ
- selected from 459 reports after several filters
- closed GitHub issues from outside contributors
- capped reports per repository
- removed trivial failures and reports developers had not reproduced
- includes 23 crash and 67 non-crash reports
- buggy app packages, scripts, configurations, and report-quality information
- at least two authors verified each retained failure
- inference: useful starting benchmark, not a random sample of all incoming mobile bugs
- selection favors failures that people could already reproduce
- Johnson et al, SANER 2022, 180-report empirical study, RQ3 and RQ4
- authors: â92.2% of the bug reports had at least one missing S2Râ
- S2R means step to reproduce
- compared initial descriptions with actions needed in manually derived scripts
- distinguished missing setup context from missing actions between reported steps
- follow-up discussions supplied relevant additional information in 35 reports
- inference: setup information can be as important as choosing the next button
- the percentage applies to this reproducible-report sample
- discussion content may include later diagnosis or fix details
- authors: â92.2% of the bug reports had at least one missing S2Râ
from text matching to global search
- Zhang et al, ISSTA 2023, ReproBot paper, method and validity discussion
- authors: âour approach is unable to reproduce non-crash bug reportsâ
- extracts actions and targets from text, then uses reinforcement learning to search for matching UI actions
- reinforcement learning updates action preferences from observed rewards
- repeated randomized runs and manual agreement on extracted-step labels
- inference: a useful non-LLM baseline for crash reproduction
- learning a path does not solve the general non-crash failure-check problem
- Zhang, Tawsif, Ryu, Yu, and Halfond, FSE 2024, Roam paper, sections 4.2â4.8
- authors: âwe provided the same step entities collected in Section 4.2 to all three approaches and Roamâ
- searches a graph of screens and actions for globally matched paths
- dynamic programming ranks sequences while accounting for missing or inaccurate steps
- retained 72 of 399 candidate reports after reproducibility and availability checks
- 46 retained reports had missing steps
- manually prepared action descriptions and crawler-built screen graphs
- reported 94% reproduction with a one-hour limit per report
- authors replayed outputs manually
- inference: strong evidence for its matching algorithm under prepared inputs
- not a demonstrated 94% fully automatic pipeline from arbitrary raw reports
- graph completeness and one-step-to-one-action assumptions limit generalization
- globally best matching path in the screen model need not reproduce the real failure when relevant state is missing
LLMs already help interpret incomplete reports
- Feng and Chen, ICSE 2024, AdbGPT paper, evaluation and threats
- authors: âwe ran the LLM-related approaches ⊠three timesâ
- ellipsis omits the parenthetical list of approach names
- prompts an LLM to extract steps and guide replay
- examples and intermediate reasoning guide interpretation
- includes a small developer study
- inference: compare against this baseline before proposing generic LLM-guided replay
- repeated runs reduce sensitivity to randomness but do not remove dataset-selection limits
- authors: âwe ran the LLM-related approaches ⊠three timesâ
- Wang et al, ISSTA 2024, ReBL paper, sections 3.5, 4.1â4.3, and 5.1
- authors: âwe conducted a manual inspection to confirm whether the described crash or non-crash bug symptom occursâ
- feeds the whole report, grouped UI information, feedback, and history to GPT-4
- model decides whether the reported symptom has appeared
- retained 96 accessible reproducible reports
- reported 69/73 crash and 18/23 non-crash reproductions
- comparisons with other tools used the crash subset
- reported mean time measures successful reproductions
- failed runs and setup costs must also enter an operational budget
- false-symptom example: cutting and pasting within the same folder produced no visible change
- reported bug concerned pasting into a different folder
- inference: non-crash recognition already exists; context-sensitive false success remains concrete
- evaluations used manual confirmation rather than treating the modelâs success declaration as ground truth
- Huang et al, TSE 2025, ReActDroid abstract
- authors: âreproduce mobile application crashes directly from the crash overviewâ
- uses reasoning, app knowledge, and exploration history to infer actions from a one-sentence overview
- limit: full evaluation not read here
- ReActDroid author artifact, dataset description
- maintainers: âFor the 59 bug reports that at least one tool can successfully reproduceâ
- describes reproduction-step and app-size statistics for this selected subset
- these statistics do not describe all reports or explain failures that every tool missed
- read-depth limit: README inspected; full-paper methods and artifact replay remain outstanding
recognizing failures is a separate research problem
- Zhang, Ryu, Yu, and Halfond, FSE 2025, BugSpot abstract
- authors: âtransforms the documented buggy behavior into this structured languageâ
- empirically classifies failure manifestations, then matches extracted descriptions against device and UI information
- inference: a proposal to create structured report-derived failure checks directly overlaps this work
- full paper could not be accessed here; no unverified accuracy figure is claimed
- BugSpot author artifact, input documentation
- maintainers: âbefore the first actionâ and âafter each UI actionâ
- expects device information, screenshots, and UI hierarchies
- before and after actions
- inference: temporal device evidence is an existing interface, not a new idea
- BugSpot parser prompt, grammar and examples
- authors: âs1 = S() representing the UI screen before the last screen, s2 = S() representing the UI screen after the last screenâ
- defines elements, screens, and device states
- element descriptions and checkbox status
- crash dialogs and keyboard visibility
- logs, volume comparisons, and audio status
- supports presence, absence, conjunction, and comparisons
- this is the released grammar, not a verified complete taxonomy from the paper
- inference: comparing before and after screens directly overlaps a proposed intermediate representation
- BugSpot recognizer implementation
- author comment: âConfidence too low. Ignore the chose widget.â
- rejects LLM element matches with confidence at most five
- this threshold is not evidence that confidence predicts correctness
- screen equality compares layout hashes
- this checks layout agreement, not successful completion of the intended task
- color and location branches contain
pass- inspection shows these filters are not implemented in this artifact version
- this does not establish what was implemented in the paperâs evaluated version
- BugSpot language-change prompt example
- example report: âI went to Setting and tried to change the language to Franch. However, the app still shows English.â
- generated check looks for English text
- it does not explicitly verify that French was selected
- inference: ordinary English output without a completed language change is a useful negative case
- test whether replay or another checker establishes the missing precondition
- no false-positive rate is claimed from reading this example
- Johnson, Mahmud, Chaparro, Moran, and Fazzini, ASE 2025, AndroB2O paper, sections III, VâVIII
- authors: âFBOs fail on the buggy appâ
- FBO means failure-based oracle
- assertions should pass on the corrected version
- generates executable UIAutomator assertions from report text and the failing screenâs XML hierarchy
- requires consistent element and assertion predictions across repeated prompts
- 152-report main evaluation reported 61.2% successful oracle generation
- separate 17-report preliminary study and 16 later reports
- benchmark provides buggy and corrected app versions
- evaluation manually checked semantic relevance as well as buggy/fixed outcomes
- unrelated concurrent changes can otherwise yield a misleading pass/fail pair
- already integrates with ReBL and studies incorrectly declared reproductions
- limitation: missing-element assertions can pass on the wrong screen
- authors suggest adding screen identity
- dataset excludes failures unavailable through UIAutomator or screen hierarchies
- inference: broad oracle generation and buggy/fixed validation have direct prior work
- extend evaluation to preconditions observable in UI action histories and matched correct outcomes before proposing a new method
- authors: âFBOs fail on the buggy appâ
- Wang et al, MSR 2025, image study, methods and RQ4
- authors: âReproduction succeeds only after manually providing the information of the imageâ
- Table III defines a result category
- sampled 367 image-containing reports for role analysis
- separately collected 42 reproducible image-containing reports for tool evaluation
- compared text-only reproduction and manual translation of image information
- images can describe steps, observed behavior, or expected behavior
- those roles require different handling
- inference: selective interpretation of report images is established prior work
- manual image translation is not demonstrated autonomous visual reasoning
- the image-containing sample does not estimate all-report reproduction success
- authors: âReproduction succeeds only after manually providing the information of the imageâ
recent planning and evaluation results narrow novelty further
- Chen et al, TreeMind preprint, February 1 2026 version, sections 4â5
- authors: âA bug is considered successfully reproduced if at least one of the five runs succeedsâ
- combines tree search with LLM-generated candidate actions and progress scores
- UI text and screenshots support action choice
- reported 60/93 crash reports reproduced
- five runs per technique, thirty-minute limit per run
- recorded shortest successful reproduction time
- inference: this is five-attempt success and selected successful latency
- it does not measure one-attempt reliability or total campaign cost
- prompting examples came from the dataset; use held-out apps when comparing new methods
- documented failures include external services and screens shown only once
- revisiting a search node need not restore the same app state
- Wang et al, ReproPilot ASE 2026 abstract
- authors: âdiverse, reusable long-horizon guidance beyond step-wise action suggestionsâ
- combines that guidance with global planning
- reports reliability and token savings on 74 crash reports
- abstract does not define consistency, repetitions, failed-run costs, or reset costs
- do not compare its headline percentages with one-run reproduction rates
- inference: caching multi-action LLM guidance and globally reprioritizing exploration already have direct prior work
- full methods needed to interpret its consistency metric and cost accounting
- Carey, Elish, and Abedin, ICST 2026 short-paper abstract
- authors: âcritical-step coverage, an intent-level metricâ
- evaluates classification and reproduction scripts against 90 AndroR2 reports
- abstract reports 70% classification accuracy and 89.5% mean critical-step coverage
- these are distinct measures
- coverage does not establish executable replay or failure reproduction
- full methods still needed for essential-step labels, annotator agreement, model settings, and execution protocol
- inference: action intent and final outcome already have separate evaluation proposals
- high essential-action coverage does not alone demonstrate the correct failure
- CARBON repository
- maintainers: âcomplex gestures like pinch-to-zoom, double-tap, edge swipesâ
- lead for current gesture-capable screenshot agents
- repository claims are not independently reproduced here
- no publication status or benchmark fairness is assumed
candidate study: near-miss checks for non-crash failures
- hypothesis: explicit checks of visible task preconditions reduce false success beyond existing history-aware judgment and generated assertions
- start with action history and visible screens
- an agent choice to keep the first comparison bounded
- start with action history and visible screens
- bounded initial scope: file movement, list editing, and saved settings
- use local open-source apps with executable buggy and corrected versions
- construct true failures and nearby correct outcomes with similar final screens
- same-folder versus different-folder paste
- empty input versus failed nonempty search
- setting changed in memory versus persisted after restart
- compare ReBLâs history-aware judgment, BugSpot, AndroB2O, and simple task-specific assertions
- hold replay sequences fixed first
- isolates checking errors from navigation errors
- then measure the effect on end-to-end search termination
- hold replay sequences fixed first
- outcomes: false acceptance, missed failures, abstentions, and checking cost
- abstention means the checker says evidence is insufficient
- independent reviewers label runs without knowing which checker produced the result
- split by app, not just reports, to limit tuning on repeated interfaces
- possible contribution: failure recognition evaluated against deliberately similar correct outcomes
- generic semantic checks are already covered by BugSpot and AndroB2O
- AndroB2O already proposes checking screen identity
- add the BugSpot language-change example as a precondition check
- ordinary English screen versus English screen after a verified French selection
- inspect BugSpotâs full taxonomy and evaluated negative cases before claiming novelty
- artifact inspection alone does not close this gap
- stop rule: abandon a new-method claim if existing assertions, BugSpot, or AndroB2O perform equally well
- an evaluation dataset may still be useful if it exposes reproducible failure-check errors
candidate study: whether search resets preserve bug-relevant state
- Li et al, ICSE-SEIP 2020, RepRev full methods, evaluation, and limitations
- authors: âcorrect event sequences, but the crash does not happenâ
- extracts informative review words and guides UI exploration by semantic similarity
- when all current actions score zero, explores one step ahead to improve their ranking
- screen abstraction ignores text-content and checkbox-state changes
- dynamic-list abstraction explores only one item in a nonempty list
- evaluation contains 63 bug-related reviews requiring manually reproduced ground truth
- authors describe three failures attributed to unknown environmental conditions
- two additional attempts still failed in those cases
- this attribution is not controlled evidence that network or hardware caused the failures
- another failure requires repeating record-and-play twice
- UI abstractions and loop-breaking rules may discard behavior needed by the reported failure
- this is an earlier concrete counterexample to equating apparent navigation progress with bug-relevant progress
- inference: compare state restoration with finer search-state representation and action-history preservation
- restoration alone cannot recover a necessary action sequence that the search policy excludes
- ReActDroid pinned environment implementation, initialization and relaunch functions
- authorsâ configuration:
desired_caps['noReset'] = True - selected Appium configuration requests retained app state and automatic permission granting
- relaunch starts an activity through adb or calls Appiumâs launch function
- these functions do not explicitly restore a snapshot or clear app data
- this is a local code observation, not an artifact-wide absence claim
- authorsâ configuration:
- ReActDroid pinned main loop, step and run functions
- authorsâ condition includes
dst_observe["page_id"] == "out of app" - leaving the app or observing an empty page triggers relaunch
- the same main object retains its search memory
- selected example stops after 100 steps
- this is not an independently verified published evaluation budget
- implication: retained app data and retained search history are separate comparison dimensions
- neither source establishes that relaunch restores all bug-relevant conditions
- reading limit: selected released code inspected without execution
- full paper and benchmark-specific runners remain unchecked
- authorsâ condition includes
- agent proposal extending beyond the UI-only method described in the defense notes
- hidden state means data or conditions not distinguishable from the current screen
- motivation from reviewed literature
- the Android reproduction study distinguishes missing setup context from missing actions
- TreeMind documents screens shown only once and dependence on external services
- inference: returning to the same-looking screen may not restore conditions needed to reproduce the same failure
- hypothesis: restoring relevant saved data and environment conditions improves repeated reproduction of stateful bugs
- controlled comparison: retained data with relaunch, restart, reinstall, emulator snapshot, and explicit app-data restoration
- use local reports requiring first-run setup or persisted edits
- hold app version, Android version, locale, permissions, and local service responses fixed
- record whether restored screens match and independently verify whether the same reported failure occurs
- separately vary whether search distinguishes persisted data and repeated action histories
- isolate reset failure from abstraction failure and missed timing requirements
- deterministic local service responses remove one alternative explanation without representing all remote-service bugs
- measure repeatable same-failure replay, reset time, storage, failed attempts, and all LLM calls
- compare total budgets rather than successful-run latency alone
- novelty remains unverified
- include ReActDroidâs retained-state relaunch baseline
- verify the relevant benchmark runner before comparing costs or results
- recover ReproPilotâs reset and history mechanisms
- its abstract does not establish whether it preserves the conditions tested here
- snapshots and state restoration are established techniques
- proposed question: which conditions must report-driven search preserve beyond screen appearance
- include ReActDroidâs retained-state relaunch baseline
- stop rule: reject a special restoration method if standard snapshots solve the cases at acceptable cost
before choosing an experiment
- retrieve BugSpotâs complete paper and inspect its evaluated negative cases
- grammar, recognizer code, and language-change example inspected
- complete taxonomy and negative-case evaluation remain unchecked
- obtain full ReActDroid, ReproPilot, and ICST intent-study methods
- method and evaluation reviews remain outstanding
- recovery attempts on October 7 2026 did not produce these full papers
- Song Wangâs author page lists ReActDroid and a PDF link
- the linked tse25.pdf returned HTTP 404
- Zhe Liuâs author page lists the paper without a download link
- Sidong Fengâs publications page labels ReproPilot âPDFâ
- the corresponding HTML link has an empty destination
- the ACM paper endpoint returned HTTP 403
- Karim Elishâs publication list identifies the ICST paper
- the author-site metadata inspected did not provide a paper download
- BugSpotâs author page and artifact did not yield its complete paper
- Song Wangâs author page lists ReActDroid and a PDF link
- these are access results in this review, not evidence that copies do not exist elsewhere
- additional October 7 routes inspected publisher text, author publication metadata, institutional records, and ResearchGate listings
- IEEE exposes the introduction but gates ReActDroidâs methods
- TUMâs ReActDroid record provides DOI metadata rather than a manuscript
- ResearchGate lists ReActDroid and BugSpot as request-full-text entries
- conference listings did not yield ReproPilotâs full methods
- new routes did not recover these three papers; RepRevâs accessible primary methods narrow the reset proposal instead
- an arXiv link on the ICST page led to a different web-testing paper
- it was excluded from this review
- inspect AndroB2Oâs released benchmark before collecting overlapping cases
- rerun available artifacts under one frozen environment and common total budget
- report one-run success and success after a fixed number of attempts separately
- audit excluded reports, unavailable app packages, and server dependencies
- complete independent review of novelty and causal comparisons
- no ChatGPT critique is claimed in this page
Last edited: