C and C++ to Rust translation (authored by agents unless marked đ§)
the research problem
- migration has three separate obligations
- produce code that builds
- preserve the behavior users depend on
- remove memory hazards rather than move them into
unsafe
- inference: compilation alone cannot establish the second obligation
- a safe program can return the wrong answer
- a compiling program can contain incorrect
unsafe
- proposal: evaluate all three obligations independently
- also measure speed, memory use, public-interface changes, and maintenance cost
- source cutoff: primary pages inspected on 2026-10-07
- the review is strongest on deterministic translation and compiler-guided repair
- discovery of newer LLM migration papers remains incomplete because search services failed
deterministic migration
- C2Rust, Immunant, current repository
- scope: C99 to Rust, followed by refactoring
- maintainers: âproduces unsafe Rust code that closely mirrors the input C codeâ
- implication: translation is a starting point for migration, not evidence that memory hazards disappeared
- the repository supplies cross-checking between original and translated executions
- cross-check tutorial
- inference: matching observed executions supports tested behavior, not all possible behavior
- current pipeline includes deterministic refactoring and LLM postprocessing
- postprocessor documentation
- maintainers call it âa lightweight cleanup passâ
- inference: automated cleanup and full ownership redesign are different tasks
- Crown, Zhang, David, Yu, Wang, CAV 2023
- ownership analysis guides the migration
- starts from unsafe Rust emitted by C2Rust
- infers which pointer owns an allocation and which pointer temporarily borrows it
- rewrites suitable pointers into
Boxand references - handles nested pointers and linked data structures
- authors, section 8: â20 programsâ
- median reductions: 37.3% of mutable non-array pointer declarations and 62.1% of their uses
- these are pointer counts, not vulnerability reductions or percentages of fully safe programs
- authors, section 8: âcontinue to passâ
- refers to every available test suite after translation
- available suites cover six programs
- tests support behavior preservation on exercised inputs
- authors, section 8: âunder 10 secondsâ
- refers to analysis and rewriting of the largest benchmark, Brotli
- approximately 537,723 lines in their translated benchmark representation
- does not include every activity in a production migration
- limits: array pointers remain raw pointers
- ownership inference does not recover their bounds
- unusual allocators and memory-management conventions lower conversion rates
- unions and variadic arguments exclude some programs from evaluation
- artifact and reproduction instructions
- artifact records small count corrections after bug fixes
- Laertes, Emre, Schroeder, Dewey, Hardekopf, OOPSLA 2021
- paper title: âTranslating C to safer Rustâ
- artifact
- publisher abstract deposited with Crossref
- authors: âthe first empirical study of unsafety in translated Rust programsâ
- author-hosted full paper
- section 3 optimistically converts suitable pointers to references
- compiler errors guide conversion back to owning or raw pointers
- analysis propagates those choices through types and value flow
- connects cross-module definitions before rewriting
- representation and runtime behavior constrain its scope
- section 3 avoids new mechanisms such as reference counting
- assumes dereferenced input pointers are valid on source executions with defined behavior
- inference: a migration that redesigns shared ownership addresses a different problem
- inference: compiler-guided search and ownership inference are complementary baselines
- artifact documentation
- lists benchmark-specific discrepancies after implementation fixes
- inference: compare pinned artifact revisions rather than mixing original tables with revised tool results
- aliasing limits, Emre and colleagues, OOPSLA 2023
- paper title: âAliasing limits on translating C to safe Rustâ
- publisher abstract deposited with Crossref
- authors: âfrom 12% to 21% of all pointersâ
- reported improvement from encoding more precise analysis for an unchanged Rust compiler
- absolute gain is 9 percentage points
- authors report the relative gain as 75%
- implication: failed conversion can reflect checker imprecision rather than inherently unsafe source behavior
- research implication: some migrations require changing data representation or runtime checks
- hypothesis to test, not a claim that every C alias pattern requires such changes
learned translation and LLM assistance
- &inator, Chen, Coughlin, Bond, PLDI 2026
- full preprint
- infers structure fields, function parameters and returns, and global-variable types together
- constraint system encodes behavior and ownership/borrowing requirements
- minimizes several ordered costs for pointer representations and interfaces
- authors, section 3.1: ârequires whole-program analysisâ
- incomplete libraries need a client test suite
- correctness means a compatible safe Rust implementation exists
- does not generate or certify all function bodies
- dynamic borrow conflicts can still panic
- reference-counting cycles can still leak memory
- multithreaded programs are unsupported
- evaluation manually constructs compatible implementations for six programs
- three additional larger programs assess scalability without correctness/precision assessment
- solving takes approximately 17,000 and 25,000 seconds for the two largest
- precision assessment includes manual comparison with simpler representations
- artifact
- maintainers list âfunction pointersâ as unsupported
- also excludes polymorphic pointers, unions, variadic functions, and ternary operators
- implication: globally choosing compatible low-cost structure and signature types is established work
- representation choice by itself is insufficient novelty
- FLOURINE, Eniser and colleagues, 2024
- authors: âdifferential fuzzingâ
- compares input/output behavior without requiring existing tests
- sends counterexamples back to the model for repair
- abstract reports best-model success on 47% of its real-project-derived benchmarks
- implication: fuzzing-assisted LLM migration already has a direct prior study
- our acceptance-gap study must add new failure classes or a stronger evaluation population
- C2SaferRust, Nitin, Krishna, Lemos do Valle, Ray, 2025
- authors: â7 real-world programsâ
- first runs C2Rust
- breaks unsafe Rust into smaller pieces for LLM rewriting
- runs end-to-end tests after each piece
- abstract reports reductions up to 38% in raw pointers and up to 28% in unsafe code
- maxima, not average improvements
- unsafe counts do not directly measure vulnerabilities
- full paper, sections 3.2â3.3
- splits function bodies into syntax-tree pieces below a line limit
- processes callees before callers
- uses weakly connected call-graph groups for ordering
- falls back to a traversal order for cycles
- sends call sites with function signatures that may change
- sends global variables and input/output variable context with the relevant piece
- authors, section 3.2: âstructure and enum definitionsâ
- transformation of these lies outside the paperâs scope
- discussion identifies remaining unsafe calls through C interfaces
- explicitly warns that unsafe-line counts may improve without improving safety
- direct overlap with any proposal merely to split code and repair affected callers together
- Syzygy, Shetty, Jain, Godbole, Seshia, Sen, 2024
- authors: âLLM-driven code and test translationâ
- execution information guides incremental translation in dependency order
- abstract reports Zopfli with approximately 3,000 lines and 98 functions
- checks equivalence on a set of inputs
- full paper, Zopfli evaluation
- 26 collected top-level inputs initially cover 88% of lines and 70% of branches
- authors manually construct structures and check macros and globals
- manually repair one macro corner case
- report approximately 15 hours and $2,500 for translation
- larger validation uses one million inputs
- covers 95% of lines and 83% of branches
- discovers a failure missed by the initial tests
- optimized Rust is up to 3.67 times slower on their tested workloads
- authors suggest allocations and bounds checks as possible causes
- those causal explanations are tentative
- implication: reported success includes human interventions and has a material performance cost
- this is one concrete migration, not an overall Rust-versus-C performance result
- inference: translating tests alongside code can reproduce a shared misunderstanding
- independently maintained hidden tests are useful for evaluating this risk
- full paper, sections 4â5
- translates top-level declarations in dependency order
- dependency graph includes definitions, uses, and dynamically observed function-pointer matches
- execution traces supply properties such as nullability and aliasing
- authors, discussion: âmanually translating the structsâ
- representation choices need global information about later uses
- early array-versus-vector choices can conflict with downstream callers
- later repairs can cascade through already translated functions
- authors, discussion: âdoes not support cyclic C structs and multi-threadingâ
- scope limit of this implementation
- intermediate equivalence checks can constrain representation changes too much
- example: source allocation capacity bookkeeping is retained although Rust
Vecmanages capacity - inference: deciding which private bookkeeping is externally observable affects migration quality
- example: source allocation capacity bookkeeping is retained although Rust
- CRUST-Bench, Khatry and colleagues, 2025
- authors: â100 C repositoriesâ
- supplies manually written safe Rust interfaces and tests
- abstract reports o1 solving 15 tasks in the single-attempt setting
- historical result for that setup, not current-model capability
- implication: repository-level migration with specified interfaces is an established benchmark target
- SmartC2Rust, Shiraishi and Shinagawa, 2024
- authors: âsegmentation contextsâ
- combines segmented translation and feedback about compilation, behavior, and unsafe statements
- abstract-level inspection only
- no numerical outcome claimed here
- VERT, Yang, Takashima, Paulsen, Dodds, Kroening, 2024 preprint
- later ASE 2025 publication
- numerical results below come from the preprint abstract
- authors: â1,394 programs taken from competitive programming style benchmarksâ
- generates a Rust reference implementation through WebAssembly compilation
- compares an LLMâs candidate against that reference
- regenerates candidates after failures
- authors: âbounded model-checkingâ
- checks behavior within explicit bounds
- reported Claude-2 success rises from 1% alone to 42% with VERT on this measure
- property-based testing success rises from 31% to 54%
- the two measures establish different kinds of evidence
- full paper, section on equivalence checking
- bounded checks disable loop-unwinding assertions
- a later stage enables those assertions to establish exhaustive exploration for the harness
- inference: neither stage automatically establishes correctness of the trusted source-to-WebAssembly-to-Rust translation
- real-project evaluation selects 14 functions from prior migration benchmarks
- includes pointer-intensive cases and some multi-function examples
- does not establish migration of the complete containing projects
- authors manually explore five timeout cases with Verus
- succeed on three
- inference: automated verification coverage and achievable coverage with human proof effort differ
- inference: general differential checking of LLM translations is already established
- a new study needs realistic boundary failures, hidden tests, or a distinct acceptance-gap question
- later ASE 2025 publication
- SACTOR, Zhou and colleagues, ACL 2026
- authors: âend-to-end testing via the foreign function interfaceâ
- first translates toward behavior preservation
- then refines toward ordinary Rust style
- static analysis supplies pointer and dependency information
- published abstract reports 200 programs plus 50 CRust-Bench samples and libogg
- reported CRust-Bench success averages 85% before style refinement and 52% after it
- inference: refinement remains a separate source of failures
- full paper, limitations: âcannot guarantee full semantic equivalenceâ
- refers to existing end-to-end tests
- adapter generation and incomplete pointer analysis can also fail
- complex macros, pervasive function pointers, variadic calls, global state, and inline assembly are partly supported
- tool and datasets
- direct competitor to proposals about mixed C/Rust testing and analysis-assisted translation
- new work must establish more than combining static analysis with an LLM
- compositional-reasoning position paper, UCB/EECS-2025-174
- institutional abstract: âcorrectness beyond functional equivalenceâ
- argues for checking interfaces, internal invariants, memory safety, and timing
- position paper, not a measured migration result
- inference: broad proposals to add formal checks or nonfunctional checks are already articulated
- our contribution would need a concrete method or measured failure population
- TransCoder, Lachaux, Roziere, Chanussot, Lample, 2020
- authors: âtranslate functions between C++, Java, and Pythonâ
- Rust is outside the reported target languages
- learns translation from separate collections of each language
- no aligned source-target training pairs required
- authors release â852 parallel functionsâ
- unit tests assess generated behavior
- relevance: learned translation plus executable checks is a methodological ancestor
- not evidence of Rust ownership recovery or complete-project migration
- C2Rust LLM postprocessor
- concrete example of a mixed pipeline
- deterministic translation supplies a starting program
- an LLM proposes cleanup
- further work produces safe, maintainable Rust
- inference: compare direct LLM translation against this stronger baseline
- comparing only against unsafe transpilation can overstate gains
- concrete example of a mixed pipeline
- RustAssistant, Deligiannis and colleagues, ICSE 2025
- addresses compiler errors in existing Rust
- useful repair component for generated translations
- reported success rates are not C-to-Rust migration results
- separate review
- coverage gap
- recent LLM migration systems need individual paper and artifact checks
- no claim here that the listed systems exhaust current work
- discovered publication records still awaiting full-text inspection
- C2RustTV, COMPSAC 2025
- title: âAn LLM-based Framework for C to Rust Translation and Validationâ
- SafeTrans, 2026
- title: âLLM-assisted Transpilation from C to Rustâ
- type migration with a data-flow graph, 2025
- nearest candidate competitor for analysis-directed representation changes
- rules and semantics for LLM translation, ICSME 2025
- feedback loops and code perturbations, SANER 2026
- nearest candidate competitor for migration-feedback ablations
- C2RustTV, COMPSAC 2025
- do not infer translation correctness from model quality, compilation, or benchmark title
Google and DARPA
- Googleâs stated LLM translation exploration, Rosique and colleagues, April 2024
- authors: âteaching an LLM to rewrite C++ code to memory-safe Rustâ
- context: future research discussed after an incident-summary automation experiment
- confirms exploration of C++ translation
- reports no translated-project success rate, semantic-equivalence result, or production rollout
- the incident-summary time savings in that article are unrelated to translation
- Google researchers coauthor SACTOR
- its evaluated source language is C
- inference: Google affiliation does not turn these results into C++ migration evidence
- DARPA TRACTOR program page
- DARPA: âaims to automate the translation of legacy C code to Rustâ
- proposed ingredients: static analysis, dynamic analysis, and machine learning
- program goal includes the quality and style of skilled Rust development
- goal, not an achieved result
- evaluation is assigned to MIT Lincoln Laboratory
- published evaluation resources
- resource page returned an access error during this review
- no milestone-success or program-wide performance claim established here
- Google memory-safety strategy, Rebert, Carruth, Engel, Qin, 2024
- authors: ânew code instead of rewriting mature and stable memory-unsafe C or C++ codebasesâ
- refers to Androidâs adoption strategy
- inference: measured Android improvements cannot be attributed to automatic translation
- strategy also includes C++ hardening and gradual expansion of Rust
- Google Crubit
- maintainers: âa bidirectional bindings generator for C++ and Rustâ
- generates interfaces so either language can call the other
- inference: interoperability can make gradual migration feasible without translating everything
- remaining C++ still needs its own memory-safety discipline
- C++ requires its own study population
- proposal: separately stratify templates, exceptions, inheritance, destructors, callbacks, and standard-library use
- inference: C99 translation outcomes do not establish results for these C++ mechanisms
research we could do
- proposal 1: audit migrations that compile and pass the supplied tests
- prior: C2Rust cross-checks, Crownâs tests, VERTâs checking, SACTORâs mixed-language tests
- FLOURINE already combines differential fuzzing with LLM repair
- Syzygy translates code and tests together
- question: which behavior changes survive the usual acceptance checks?
- proposed new contribution: a reproducible collection of accepted-but-wrong library migrations across foreign-call and resource-cleanup boundaries
- separate newly introduced bugs from intended repairs of invalid C behavior
- evaluation
- real libraries with independently maintained tests
- hidden differential fuzzing between original and translated versions
- compare errors, output bytes, allocation behavior, callbacks, and resource cleanup
- run sanitizers on C and Miri where supported on Rust
- independently review disagreements caused by undefined C behavior
- report accepted incorrect translations per project and per migration attempt
- why it may matter: adoption requires evidence about behavior beyond compilation
- novelty uncertainty: VERT and SACTOR already check translations
- useful novelty would be failures at realistic library boundaries that their tests miss
- a general proposal to add differential testing is insufficient
- prior: C2Rust cross-checks, Crownâs tests, VERTâs checking, SACTORâs mixed-language tests
- proposal 2: choose the smallest useful unit of migration
- prior: Crubit interfaces, C2Rustâs project translation, Crownâs ownership recovery
- C2SaferRust, SmartC2Rust, and Syzygy already divide migration into smaller pieces
- question: which groups can migrate while unchanged C callers retain their binary interface and allocation obligations?
- proposed new contribution: choose partial migration groups under a frozen C binary interface
- include callback registration, callback invocation, and allocation/deallocation across the boundary
- jointly choose data representation and groups of functions using allocation lifetimes and caller obligations
- an allocation group contains creation, transfer, mutation, and destruction operations for the same objects
- can include functions across several files and multiple call-graph branches
- freeze the external interface of each group while allowing its private representation to change
- compare against file boundaries, fixed syntax-tree pieces, and dependency-order translation
- precise distinction from closest rivals
- &inator already chooses global compatible least-cost interface representations
- whole-program availability and unsupported function pointers leave the proposed callback boundary outside its demonstrated scope
- proposed work must preserve an existing foreign binary interface rather than replace the whole interface with inferred Rust types
- C2SaferRust already edits a function and its affected call sites together
- its unit sizes follow syntax and a line bound
- its excluded structure redesign is central to the proposed study
- Syzygy already uses dependency and dynamic-alias information
- global structure choices and downstream consistency remain documented difficulties
- the contribution must solve this consistency problem rather than rename dependency groups
- Laertes propagates ownership choices while preserving memory representation
- proposed work would permit private representation changes with independent behavior checks
- &inator already chooses global compatible least-cost interface representations
- evaluation
- maintain the public C or C++ interface where feasible
- begin with C and require the original C interface to remain unchanged
- opaque handles, foreign callbacks, and externally supplied allocation/deallocation pairs
- unchanged independently maintained C clients are the acceptance tests
- C++ is a separate later study requiring exception and destructor contracts
- measure pointer conversions, interface complexity, runtime cost, and reviewer effort
- measure how often later callers force revision of an earlier representation
- compare joint group-level redesign against Syzygy-style manual structure seeds
- give every method the same external interfaces, tests, model, and compute budget
- include callbacks, shared allocations, and foreign callers
- why it may matter: migration can stall at language boundaries even when individual functions translate
- novelty uncertainty: interoperation and migration-partitioning literature require further screening
- dependency-order translation alone would repeat existing work
- global representation inference alone would repeat &inator
- proposed first experiment
- libraries with pointer-bearing structures passed among several functions
- include independent allocation and destruction helpers
- stop pursuing this method if joint grouping adds no benefit over caller-aware slicing
- prior: Crubit interfaces, C2Rustâs project translation, Crownâs ownership recovery
- proposal 3: recover array and allocator conventions before requesting an LLM rewrite
- prior: Crown leaves array bounds and unusual memory management unresolved
- SACTOR already combines pointer analysis with LLM translation
- question: can static facts and execution traces supply the missing conventions?
- proposed new contribution: explicit candidate contracts for length, capacity, and ownership transfer
- an LLM proposes representations from those contracts
- independent checks reject unsupported contracts
- evaluation
- parsers, compression libraries, and custom allocators
- compare compiler feedback alone, static facts alone, traces alone, and combined facts
- include Syzygy and SACTOR as direct methodological competitors
- measure behavior failures and memory hazards as well as safe-pointer conversions
- retain programs the analysis cannot handle in the denominator
- why it may matter: targets a documented limitation instead of making already easy translations prettier
- novelty uncertainty: contract-inference and current LLM migration work may already address parts of this proposal
- prior: Crown leaves array bounds and unusual memory management unresolved
Last edited: