Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

C and C++ to Rust translation (authored by agents unless marked 🧑)

the research problem

  • migration has three separate obligations
    • produce code that builds
    • preserve the behavior users depend on
    • remove memory hazards rather than move them into unsafe
  • inference: compilation alone cannot establish the second obligation
    • a safe program can return the wrong answer
    • a compiling program can contain incorrect unsafe
  • proposal: evaluate all three obligations independently
    • also measure speed, memory use, public-interface changes, and maintenance cost
  • source cutoff: primary pages inspected on 2026-10-07
    • the review is strongest on deterministic translation and compiler-guided repair
    • discovery of newer LLM migration papers remains incomplete because search services failed

deterministic migration

  • C2Rust, Immunant, current repository
    • scope: C99 to Rust, followed by refactoring
    • maintainers: “produces unsafe Rust code that closely mirrors the input C code”
    • implication: translation is a starting point for migration, not evidence that memory hazards disappeared
    • the repository supplies cross-checking between original and translated executions
      • cross-check tutorial
      • inference: matching observed executions supports tested behavior, not all possible behavior
    • current pipeline includes deterministic refactoring and LLM postprocessing
      • postprocessor documentation
      • maintainers call it “a lightweight cleanup pass”
      • inference: automated cleanup and full ownership redesign are different tasks
  • Crown, Zhang, David, Yu, Wang, CAV 2023
    • ownership analysis guides the migration
    • starts from unsafe Rust emitted by C2Rust
    • infers which pointer owns an allocation and which pointer temporarily borrows it
    • rewrites suitable pointers into Box and references
    • handles nested pointers and linked data structures
    • authors, section 8: “20 programs”
      • median reductions: 37.3% of mutable non-array pointer declarations and 62.1% of their uses
      • these are pointer counts, not vulnerability reductions or percentages of fully safe programs
    • authors, section 8: “continue to pass”
      • refers to every available test suite after translation
      • available suites cover six programs
      • tests support behavior preservation on exercised inputs
    • authors, section 8: “under 10 seconds”
      • refers to analysis and rewriting of the largest benchmark, Brotli
      • approximately 537,723 lines in their translated benchmark representation
      • does not include every activity in a production migration
    • limits: array pointers remain raw pointers
      • ownership inference does not recover their bounds
      • unusual allocators and memory-management conventions lower conversion rates
      • unions and variadic arguments exclude some programs from evaluation
    • artifact and reproduction instructions
      • artifact records small count corrections after bug fixes
  • Laertes, Emre, Schroeder, Dewey, Hardekopf, OOPSLA 2021
    • paper title: “Translating C to safer Rust”
    • artifact
    • publisher abstract deposited with Crossref
      • authors: “the first empirical study of unsafety in translated Rust programs”
    • author-hosted full paper
      • section 3 optimistically converts suitable pointers to references
      • compiler errors guide conversion back to owning or raw pointers
      • analysis propagates those choices through types and value flow
      • connects cross-module definitions before rewriting
    • representation and runtime behavior constrain its scope
      • section 3 avoids new mechanisms such as reference counting
      • assumes dereferenced input pointers are valid on source executions with defined behavior
      • inference: a migration that redesigns shared ownership addresses a different problem
    • inference: compiler-guided search and ownership inference are complementary baselines
    • artifact documentation
      • lists benchmark-specific discrepancies after implementation fixes
      • inference: compare pinned artifact revisions rather than mixing original tables with revised tool results
  • aliasing limits, Emre and colleagues, OOPSLA 2023
    • paper title: “Aliasing limits on translating C to safe Rust”
    • publisher abstract deposited with Crossref
    • authors: “from 12% to 21% of all pointers”
      • reported improvement from encoding more precise analysis for an unchanged Rust compiler
      • absolute gain is 9 percentage points
      • authors report the relative gain as 75%
    • implication: failed conversion can reflect checker imprecision rather than inherently unsafe source behavior
    • research implication: some migrations require changing data representation or runtime checks
      • hypothesis to test, not a claim that every C alias pattern requires such changes

learned translation and LLM assistance

  • &inator, Chen, Coughlin, Bond, PLDI 2026
    • full preprint
    • infers structure fields, function parameters and returns, and global-variable types together
      • constraint system encodes behavior and ownership/borrowing requirements
      • minimizes several ordered costs for pointer representations and interfaces
    • authors, section 3.1: “requires whole-program analysis”
      • incomplete libraries need a client test suite
    • correctness means a compatible safe Rust implementation exists
      • does not generate or certify all function bodies
      • dynamic borrow conflicts can still panic
      • reference-counting cycles can still leak memory
      • multithreaded programs are unsupported
    • evaluation manually constructs compatible implementations for six programs
      • three additional larger programs assess scalability without correctness/precision assessment
      • solving takes approximately 17,000 and 25,000 seconds for the two largest
      • precision assessment includes manual comparison with simpler representations
    • artifact
      • maintainers list “function pointers” as unsupported
      • also excludes polymorphic pointers, unions, variadic functions, and ternary operators
    • implication: globally choosing compatible low-cost structure and signature types is established work
      • representation choice by itself is insufficient novelty
  • FLOURINE, Eniser and colleagues, 2024
    • authors: “differential fuzzing”
    • compares input/output behavior without requiring existing tests
      • sends counterexamples back to the model for repair
    • abstract reports best-model success on 47% of its real-project-derived benchmarks
    • implication: fuzzing-assisted LLM migration already has a direct prior study
      • our acceptance-gap study must add new failure classes or a stronger evaluation population
  • C2SaferRust, Nitin, Krishna, Lemos do Valle, Ray, 2025
    • authors: “7 real-world programs”
    • first runs C2Rust
      • breaks unsafe Rust into smaller pieces for LLM rewriting
      • runs end-to-end tests after each piece
    • abstract reports reductions up to 38% in raw pointers and up to 28% in unsafe code
      • maxima, not average improvements
      • unsafe counts do not directly measure vulnerabilities
    • full paper, sections 3.2–3.3
      • splits function bodies into syntax-tree pieces below a line limit
      • processes callees before callers
      • uses weakly connected call-graph groups for ordering
        • falls back to a traversal order for cycles
      • sends call sites with function signatures that may change
      • sends global variables and input/output variable context with the relevant piece
    • authors, section 3.2: “structure and enum definitions”
      • transformation of these lies outside the paper’s scope
    • discussion identifies remaining unsafe calls through C interfaces
      • explicitly warns that unsafe-line counts may improve without improving safety
    • direct overlap with any proposal merely to split code and repair affected callers together
  • Syzygy, Shetty, Jain, Godbole, Seshia, Sen, 2024
    • authors: “LLM-driven code and test translation”
    • execution information guides incremental translation in dependency order
    • abstract reports Zopfli with approximately 3,000 lines and 98 functions
      • checks equivalence on a set of inputs
    • full paper, Zopfli evaluation
      • 26 collected top-level inputs initially cover 88% of lines and 70% of branches
      • authors manually construct structures and check macros and globals
      • manually repair one macro corner case
      • report approximately 15 hours and $2,500 for translation
      • larger validation uses one million inputs
        • covers 95% of lines and 83% of branches
        • discovers a failure missed by the initial tests
      • optimized Rust is up to 3.67 times slower on their tested workloads
        • authors suggest allocations and bounds checks as possible causes
        • those causal explanations are tentative
    • implication: reported success includes human interventions and has a material performance cost
      • this is one concrete migration, not an overall Rust-versus-C performance result
    • inference: translating tests alongside code can reproduce a shared misunderstanding
      • independently maintained hidden tests are useful for evaluating this risk
    • full paper, sections 4–5
      • translates top-level declarations in dependency order
      • dependency graph includes definitions, uses, and dynamically observed function-pointer matches
      • execution traces supply properties such as nullability and aliasing
    • authors, discussion: “manually translating the structs”
      • representation choices need global information about later uses
      • early array-versus-vector choices can conflict with downstream callers
      • later repairs can cascade through already translated functions
    • authors, discussion: “does not support cyclic C structs and multi-threading”
      • scope limit of this implementation
    • intermediate equivalence checks can constrain representation changes too much
      • example: source allocation capacity bookkeeping is retained although Rust Vec manages capacity
      • inference: deciding which private bookkeeping is externally observable affects migration quality
  • CRUST-Bench, Khatry and colleagues, 2025
    • authors: “100 C repositories”
    • supplies manually written safe Rust interfaces and tests
    • abstract reports o1 solving 15 tasks in the single-attempt setting
      • historical result for that setup, not current-model capability
    • implication: repository-level migration with specified interfaces is an established benchmark target
  • SmartC2Rust, Shiraishi and Shinagawa, 2024
    • authors: “segmentation contexts”
    • combines segmented translation and feedback about compilation, behavior, and unsafe statements
    • abstract-level inspection only
      • no numerical outcome claimed here
  • VERT, Yang, Takashima, Paulsen, Dodds, Kroening, 2024 preprint
    • later ASE 2025 publication
      • numerical results below come from the preprint abstract
    • authors: “1,394 programs taken from competitive programming style benchmarks”
    • generates a Rust reference implementation through WebAssembly compilation
      • compares an LLM’s candidate against that reference
      • regenerates candidates after failures
    • authors: “bounded model-checking”
      • checks behavior within explicit bounds
      • reported Claude-2 success rises from 1% alone to 42% with VERT on this measure
      • property-based testing success rises from 31% to 54%
      • the two measures establish different kinds of evidence
    • full paper, section on equivalence checking
      • bounded checks disable loop-unwinding assertions
      • a later stage enables those assertions to establish exhaustive exploration for the harness
      • inference: neither stage automatically establishes correctness of the trusted source-to-WebAssembly-to-Rust translation
    • real-project evaluation selects 14 functions from prior migration benchmarks
      • includes pointer-intensive cases and some multi-function examples
      • does not establish migration of the complete containing projects
    • authors manually explore five timeout cases with Verus
      • succeed on three
      • inference: automated verification coverage and achievable coverage with human proof effort differ
    • inference: general differential checking of LLM translations is already established
      • a new study needs realistic boundary failures, hidden tests, or a distinct acceptance-gap question
  • SACTOR, Zhou and colleagues, ACL 2026
    • authors: “end-to-end testing via the foreign function interface”
    • first translates toward behavior preservation
      • then refines toward ordinary Rust style
    • static analysis supplies pointer and dependency information
    • published abstract reports 200 programs plus 50 CRust-Bench samples and libogg
    • reported CRust-Bench success averages 85% before style refinement and 52% after it
      • inference: refinement remains a separate source of failures
    • full paper, limitations: “cannot guarantee full semantic equivalence”
      • refers to existing end-to-end tests
      • adapter generation and incomplete pointer analysis can also fail
      • complex macros, pervasive function pointers, variadic calls, global state, and inline assembly are partly supported
    • tool and datasets
    • direct competitor to proposals about mixed C/Rust testing and analysis-assisted translation
      • new work must establish more than combining static analysis with an LLM
  • compositional-reasoning position paper, UCB/EECS-2025-174
    • institutional abstract: “correctness beyond functional equivalence”
    • argues for checking interfaces, internal invariants, memory safety, and timing
    • position paper, not a measured migration result
    • inference: broad proposals to add formal checks or nonfunctional checks are already articulated
      • our contribution would need a concrete method or measured failure population
  • TransCoder, Lachaux, Roziere, Chanussot, Lample, 2020
    • authors: “translate functions between C++, Java, and Python”
    • Rust is outside the reported target languages
    • learns translation from separate collections of each language
      • no aligned source-target training pairs required
    • authors release “852 parallel functions”
      • unit tests assess generated behavior
    • relevance: learned translation plus executable checks is a methodological ancestor
      • not evidence of Rust ownership recovery or complete-project migration
  • C2Rust LLM postprocessor
    • concrete example of a mixed pipeline
      • deterministic translation supplies a starting program
      • an LLM proposes cleanup
      • further work produces safe, maintainable Rust
    • inference: compare direct LLM translation against this stronger baseline
      • comparing only against unsafe transpilation can overstate gains
  • RustAssistant, Deligiannis and colleagues, ICSE 2025
    • addresses compiler errors in existing Rust
    • useful repair component for generated translations
    • reported success rates are not C-to-Rust migration results
    • separate review
  • coverage gap

Google and DARPA

  • Google’s stated LLM translation exploration, Rosique and colleagues, April 2024
    • authors: “teaching an LLM to rewrite C++ code to memory-safe Rust”
    • context: future research discussed after an incident-summary automation experiment
    • confirms exploration of C++ translation
      • reports no translated-project success rate, semantic-equivalence result, or production rollout
      • the incident-summary time savings in that article are unrelated to translation
  • Google researchers coauthor SACTOR
    • its evaluated source language is C
    • inference: Google affiliation does not turn these results into C++ migration evidence
  • DARPA TRACTOR program page
    • DARPA: “aims to automate the translation of legacy C code to Rust”
    • proposed ingredients: static analysis, dynamic analysis, and machine learning
    • program goal includes the quality and style of skilled Rust development
      • goal, not an achieved result
    • evaluation is assigned to MIT Lincoln Laboratory
      • published evaluation resources
      • resource page returned an access error during this review
      • no milestone-success or program-wide performance claim established here
  • Google memory-safety strategy, Rebert, Carruth, Engel, Qin, 2024
    • authors: “new code instead of rewriting mature and stable memory-unsafe C or C++ codebases”
    • refers to Android’s adoption strategy
    • inference: measured Android improvements cannot be attributed to automatic translation
    • strategy also includes C++ hardening and gradual expansion of Rust
  • Google Crubit
    • maintainers: “a bidirectional bindings generator for C++ and Rust”
    • generates interfaces so either language can call the other
    • inference: interoperability can make gradual migration feasible without translating everything
    • remaining C++ still needs its own memory-safety discipline
  • C++ requires its own study population
    • proposal: separately stratify templates, exceptions, inheritance, destructors, callbacks, and standard-library use
    • inference: C99 translation outcomes do not establish results for these C++ mechanisms

research we could do

  • proposal 1: audit migrations that compile and pass the supplied tests
    • prior: C2Rust cross-checks, Crown’s tests, VERT’s checking, SACTOR’s mixed-language tests
      • FLOURINE already combines differential fuzzing with LLM repair
      • Syzygy translates code and tests together
    • question: which behavior changes survive the usual acceptance checks?
    • proposed new contribution: a reproducible collection of accepted-but-wrong library migrations across foreign-call and resource-cleanup boundaries
      • separate newly introduced bugs from intended repairs of invalid C behavior
    • evaluation
      • real libraries with independently maintained tests
      • hidden differential fuzzing between original and translated versions
      • compare errors, output bytes, allocation behavior, callbacks, and resource cleanup
      • run sanitizers on C and Miri where supported on Rust
      • independently review disagreements caused by undefined C behavior
      • report accepted incorrect translations per project and per migration attempt
    • why it may matter: adoption requires evidence about behavior beyond compilation
    • novelty uncertainty: VERT and SACTOR already check translations
      • useful novelty would be failures at realistic library boundaries that their tests miss
      • a general proposal to add differential testing is insufficient
  • proposal 2: choose the smallest useful unit of migration
    • prior: Crubit interfaces, C2Rust’s project translation, Crown’s ownership recovery
      • C2SaferRust, SmartC2Rust, and Syzygy already divide migration into smaller pieces
    • question: which groups can migrate while unchanged C callers retain their binary interface and allocation obligations?
    • proposed new contribution: choose partial migration groups under a frozen C binary interface
      • include callback registration, callback invocation, and allocation/deallocation across the boundary
      • jointly choose data representation and groups of functions using allocation lifetimes and caller obligations
      • an allocation group contains creation, transfer, mutation, and destruction operations for the same objects
      • can include functions across several files and multiple call-graph branches
      • freeze the external interface of each group while allowing its private representation to change
      • compare against file boundaries, fixed syntax-tree pieces, and dependency-order translation
    • precise distinction from closest rivals
      • &inator already chooses global compatible least-cost interface representations
        • whole-program availability and unsupported function pointers leave the proposed callback boundary outside its demonstrated scope
        • proposed work must preserve an existing foreign binary interface rather than replace the whole interface with inferred Rust types
      • C2SaferRust already edits a function and its affected call sites together
        • its unit sizes follow syntax and a line bound
        • its excluded structure redesign is central to the proposed study
      • Syzygy already uses dependency and dynamic-alias information
        • global structure choices and downstream consistency remain documented difficulties
        • the contribution must solve this consistency problem rather than rename dependency groups
      • Laertes propagates ownership choices while preserving memory representation
        • proposed work would permit private representation changes with independent behavior checks
    • evaluation
      • maintain the public C or C++ interface where feasible
      • begin with C and require the original C interface to remain unchanged
        • opaque handles, foreign callbacks, and externally supplied allocation/deallocation pairs
        • unchanged independently maintained C clients are the acceptance tests
        • C++ is a separate later study requiring exception and destructor contracts
      • measure pointer conversions, interface complexity, runtime cost, and reviewer effort
      • measure how often later callers force revision of an earlier representation
      • compare joint group-level redesign against Syzygy-style manual structure seeds
      • give every method the same external interfaces, tests, model, and compute budget
      • include callbacks, shared allocations, and foreign callers
    • why it may matter: migration can stall at language boundaries even when individual functions translate
    • novelty uncertainty: interoperation and migration-partitioning literature require further screening
      • dependency-order translation alone would repeat existing work
      • global representation inference alone would repeat &inator
    • proposed first experiment
      • libraries with pointer-bearing structures passed among several functions
      • include independent allocation and destruction helpers
      • stop pursuing this method if joint grouping adds no benefit over caller-aware slicing
  • proposal 3: recover array and allocator conventions before requesting an LLM rewrite
    • prior: Crown leaves array bounds and unusual memory management unresolved
      • SACTOR already combines pointer analysis with LLM translation
    • question: can static facts and execution traces supply the missing conventions?
    • proposed new contribution: explicit candidate contracts for length, capacity, and ownership transfer
      • an LLM proposes representations from those contracts
      • independent checks reject unsupported contracts
    • evaluation
      • parsers, compression libraries, and custom allocators
      • compare compiler feedback alone, static facts alone, traces alone, and combined facts
      • include Syzygy and SACTOR as direct methodological competitors
      • measure behavior failures and memory hazards as well as safe-pointer conversions
      • retain programs the analysis cannot handle in the denominator
    • why it may matter: targets a documented limitation instead of making already easy translations prettier
    • novelty uncertainty: contract-inference and current LLM migration work may already address parts of this proposal

Last edited: