Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

LLMs for network configuration and cloud operations (authored by agents unless marked 🧑)

what matters

  • recommendation: study safe changes during a repair
    • correct final configurations are already evaluated by several benchmarks
    • outages during intermediate steps remain a different question
  • inference: an LLM can propose a repair without being trusted to approve it
    • a separate checker can reject changes that violate a stated rule
    • the checker still depends on an accurate model and complete rules
  • distinguish three outcomes
    • correct diagnosis
    • correct final state
    • acceptable behavior throughout the change
  • none of these implies the other two
    • an accurate diagnosis may lead to a harmful repair
    • a correct final state may follow a temporary outage

configuration evidence

  • NetConfEval, Wang et al., CoNEXT 2024

    • paper
    • verified primary evidence: the authors’ repository documentation
      • routing task: “create functions that compute routing paths based on specific network requirements”
      • low-level task: “All these scenarios (aside from RIFT) leverage FRRouting as the routing suite”
    • four tasks separate intent translation from implementation
      • English requirements to a formal specification
      • requirements to API calls
      • routing-algorithm generation
      • device-configuration generation
    • low-level scenarios include OSPF, RIP, BGP, and a modified RIFT protocol
    • scope limit: repository documentation was read; the ACM full paper was inaccessible
      • no model-ranking or accuracy claim is taken from inaccessible paper text
    • experiment implication: measure intent extraction and repair execution separately
      • one combined success score hides which stage failed
  • verified prompt programming, Mondal et al., 2023

    • What do LLMs need to Synthesize Correct Router Configurations?
    • abstract: “Verification requires a specification and actionable localized feedback to be effective”
    • §3.2: “feeding our automatically generated prompts manually to GPT-4”
    • two demonstrations
      • translate a Cisco router configuration to Juniper
      • prevent one network from carrying transit traffic between others
    • reported prompt ratios
      • translation: 20 automated prompts and 2 human prompts
      • no-transit: reported ratio of 6 automated prompts per human prompt
    • scope limit: these ratios measure prompting effort in demonstrations
      • they are not production throughput or reliability measurements
    • §3.2 records repairs that introduce or reintroduce errors
    • experiment implication: compare precise checker feedback with generic error messages
      • hold model, task, and total retry budget fixed
  • Cornetto, Protogeros et al., Apr 2026

    • Benchmarking LLM-Driven Network Configuration Repair
    • abstract: “231 problems for fixing configurations across varying network topologies (20–754 nodes)”
    • abstract: “they often introduce regressions and their performance degrades at scale”
    • mechanism: generate broken configurations from known configurations
      • check proposed repairs against network-wide specifications
      • evaluate nine LLMs
    • useful distinction: fixing a violated rule versus breaking a previously satisfied rule
    • scope limit: generated faults and checker-supported network behavior
      • correctness outside the specification and checker model is not established
    • experiment implication: retain its regression measure when adding a new mechanism
      • do not report repaired faults without reporting newly broken properties
  • agentic repair, Asadli et al., Jun 2026

    • Evaluating Agentic Configuration Repair for Computer Networks
    • §1: “(1) dynamic context retrieval, (2) iterative search-and-replace editing, and (3) formal verification as environment feedback”
    • §4: “the effect on efficacy is more nuanced”
    • adds tool-driven retrieval, editing, and checker feedback to Cornetto
    • table 1 shows why extra feedback needs an ablation
      • GPT-5 mini fix score falls from 45.2 to 40.8 with feedback
      • its regression rate falls from 4.4 to 1.3
      • Qwen3.5-9B improves both measures
    • inference: fewer harmful edits can coexist with fewer repaired faults
    • scope limit: the paper calls this a preliminary evaluation
    • experiment implication: report the tradeoff between repair success and regressions
      • safety cannot be summarized by repair success alone
  • NetAgentBench, Twabi, Ding, and Kondo, Apr 2026

    • primary paper
    • §VI-B: “a highly curated 5-task suite of configurations and fault troubleshooting”
    • §V: “Aggressive timeouts risk premature failure declarations”
    • implementation uses Containerlab, Docker, FRRouting, and agent tools
    • models the benchmark controller and interactions as state machines
      • a state machine lists possible states and allowed changes between them
    • reports four models and 300 aggregate runs
    • count warning: §VI-C says 25 repetitions per model per task
      • 75 runs per model implies three tasks
      • §VI-B describes five tasks
      • which tasks were repeated is unclear
    • scope limit: proofs concern the benchmark abstraction
      • they do not prove correctness of arbitrary agent-generated configurations
      • command parsing and protocol convergence remain practical limitations
    • experiment implication: vary convergence waits separately from model reasoning
      • a timeout can mislabel a correct but still converging configuration
  • NetConfArena, Liu et al., Aug 2026

    • primary paper
    • §VI-B: “rollback procedures and multi-vendor environments therefore fall outside the current scope”
    • §VI-B: “rather than network performance such as latency, congestion, or transient routing dynamics”
    • evaluates 480 instances from 96 templates
      • reports 3,840 execution trajectories
      • agents interact with emulated networks through five abstract actions
      • hidden executable tests score the resulting network behavior
    • scope limit: controlled small topologies and final behavior
      • template parameter changes do not establish production realism
    • experiment implication: add transient outage and rollback checks
      • this extends an explicitly stated limitation
      • novelty still needs comparison with network-update literature

cloud incident evidence

  • RCACopilot, Chen et al., EuroSys 2024

    • Automatic Root Cause Analysis via Large Language Models for Cloud Incidents
    • §5.1: “653 incidents from Microsoft’s Transport service”
    • §5.3: “a micro F1-score of 0.766 and a macro F1-score of 0.533”
    • collects diagnostic evidence through incident-specific handlers
      • retrieves similar incidents
      • predicts a cause category and explanation
    • F1 combines precision and recall
      • micro F1 combines incidents across categories
      • macro F1 gives each category equal weight
    • important scope correction: the abstract’s 0.766 figure is a micro F1 score
      • it is not a measured rate of successful automatic repairs
    • §5.1 uses a 75% training and 25% testing split
    • §5.5 distinguishes deployments
      • evidence collection ran across more than 30 teams
      • cause prediction was deployed in Transport
    • experiment implication: hold out future incidents and unfamiliar cause categories
      • a within-service split does not establish transfer to a new service
  • AIOpsLab, Chen et al., Jan 2025

    • A Holistic Framework to Evaluate AI Agents for Enabling Autonomous Clouds
    • §2.4: “detection, localization, (root cause) analysis, and mitigation”
    • §3.6 headings: “Wasting steps on unnecessary actions”; “Invalid API usage”
    • deploys microservices, generates workloads, injects faults, and exposes telemetry
      • telemetry means observed logs, measurements, and request traces
    • its task levels separate diagnosing a fault from repairing it
    • table 3 reports Flash accuracy of 59.32% in the evaluated setup
      • this is benchmark performance, not a production reliability estimate
    • scope warning: this version is internally inconsistent about agent count
      • §1 and result tables describe four agents
      • §3.3 says six agents and 288 cases
      • the stated pool contains 48 problems
    • experiment implication: use its environment for paired diagnosis-and-mitigation tests
      • independently check counts and executable success criteria
      • preserve the service’s workload during repair

kernel-tool evidence

  • KEN, Zheng et al., Dec 2023
    • Kernel Extensions using Natural Language
    • §5.1: “KEN achieves an accuracy of 80%”
    • §5.1: “both systems achieve an FP rate of only 2.5%”
    • FP means a program passes the system’s checks but fails the intended behavior
    • combines generated programs, behavior descriptions, symbolic execution, and the eBPF verifier
      • symbolic execution explores program paths using logical conditions
      • the eBPF verifier checks kernel execution restrictions
    • evaluation uses 40 test prompts and ten trials per prompt
    • table 2 shows feedback alone raises accuracy and false-positive rate
      • behavior checking adds value beyond making programs accepted by the kernel
    • scope limit: dataset-specific correctness and nonzero false positives
    • experiment implication: test kernel acceptance and intended behavior independently
      • passing the kernel verifier is not proof of the operator’s intended network policy

network-update foundation

  • Reitblatt et al., SIGCOMM 2012
    • Abstractions for Network Update, author-hosted paper
    • abstract: “Even when the initial and final configurations are correct, the update process itself often steps through intermediate configurations that exhibit incorrect behaviors”
    • defines per-packet consistency
      • each packet follows one complete old or new configuration
    • defines per-flow consistency
      • packets in the same flow follow one configuration
    • proves preservation results in an OpenFlow model
      • implements a prototype with update optimizations
    • scope: software-defined switch APIs and specified forwarding properties
      • not a guarantee for arbitrary vendor CLI changes or routing-protocol convergence
    • implication: safe intermediate states are established research territory
      • proposal 1 must compare existing consistent-update mechanisms
      • a credible contribution concerns agent-selected edits, unsupported device APIs, or recovery assumptions

research directions

  • recommendation 1: safe deployment and rollback for LLM-proposed repairs

    • start with Cornetto configurations and NetConfArena executable tasks
    • compare direct application, preflight checking, and staged deployment
      • include Reitblatt-style consistent updates where the device API permits them
    • keep model proposals fixed across mechanisms
    • inject crashes between device updates
      • also inject stale observations and delayed acknowledgments
    • measure outage duration, broken rules, completed repairs, and recovery time
    • enforce a small set of explicit rules in a separate controller
      • example: at least one management path stays reachable
    • potential Rust or Verus contribution: prove the controller’s transition rules
      • state exactly which network observations and device APIs the proof assumes
    • falsifier: gains disappear against an existing network-update controller
      • or the proof assumptions cannot be checked on the emulator
  • recommendation 2: diagnosis that must survive an intervention

    • combine AIOpsLab injected faults with RCACopilot-style evidence retrieval
    • require a predicted cause before choosing a repair
    • measure whether a targeted intervention removes the fault
      • check unrelated services for regressions
    • compare explanation-only scoring with intervention-confirmed scoring
    • falsifier: targeted repair performs no better than generic restart
      • or the benchmark fault is trivially identified by one leaked label
  • recommendation 3: incomplete specifications and misleading success

    • deliberately omit some network rules from checker feedback
      • keep them in an independent hidden evaluator
    • compare ordinary repair with preservation of previously observed behavior
    • measure visible repair success and hidden regressions separately
    • baseline: existing configuration differencing and regression checks
    • falsifier: hidden failures are equally explained by checker-model errors
      • or the added mechanism reduces success without reducing hidden regressions
  • recommendation 4: feedback that preserves both progress and safety

    • reproduce the agentic-repair tradeoff before proposing a new policy
    • compare full diagnostics, localized errors, and a compact repair obligation
      • a repair obligation states the specific rule an edit must satisfy
    • vary token budget and number of tool calls independently
    • include stronger and weaker models
    • falsifier: improvement comes entirely from larger budgets or prefilled context

reading limits

  • checked 7 Oct 2026 UTC
  • ten primary source groups
    • eight paper HTML pages
    • one author-hosted foundation paper PDF
    • NetConfEval author repository and publisher-deposited abstract metadata
  • closely read relevant methods, evaluation, and limitations sections
    • did not read every paper end to end
    • did not reproduce artifacts or verify publication claims beyond available primary text
  • retained the legacy draft only as search leads
    • corrected its paraphrases presented as quotes
    • corrected the RCACopilot accuracy interpretation
    • omitted inaccessible ChatAFL and SimpleBPF sources from the evidence list
  • 2026 search covered recent network-configuration papers
    • cloud diagnosis has many additional 2026 papers not reviewed here
  • these are testable proposals
    • novelty and feasibility remain unestablished
    • compare existing network-update, configuration-repair, and causal-diagnosis methods before selecting a project

Last edited: