الملخص

Determining whether two programs are functionally equivalent is central to code modernization, patch validation, refactoring, and code-generation evaluation. Yet the usual signals are incomplete: tests cover only finite inputs, textual similarity confuses implementation with behavior, and unconstrained LLM judgments are difficult to audit. Direct execution is often impossible when a program depends on an obsolete, licensed, unavailable, or unsafe environment. We introduce FEAgent, a selective equivalence assessor agent that combines typed program-graph evidence with differential surrogate execution. FEAgent first aligns public interfaces and behaviorally relevant graph anchors, then issues bounded queries over call-flow, control-flow, data-flow, type, import, and effect relations. Next, a branch-aware generator agent proposes discriminating inputs, and two blinded LLM surrogates independently predict source and target observables. Every claim and predicted divergence is recorded in an evidence ledger. A deterministic reconciler then returns EQUIVALENT, INEQUIVALENT, or UNCLEAR rather than forcing a verdict when paths are uncovered or evidence conflicts. We evaluate FEAgent on function-level equivalence and repository-level bug patches, where the existing oracle is a benchmark label or a passing test suite. Every disagreement with that oracle is adjudicated by direct execution, revealing errors in benchmark labels and behavioral divergences missed by unit-test-only scoring. On EquiBench, execution confirms FEAgent's disagreements with published labels on 216 of 1,200 evaluated pairs (18.0%); on SWE-bench Verified, 94 of 331 test-passing agent patches (28.4%) diverge from the reference patch. FEAgent thus serves as an audit layer between testing and formal verification, keeping its evidence reviewable and its uncertainty explicit without claiming a proof of equivalence.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Kachroo, A., Hui, L., Mao, H., Zhang, Y., & Vo, N. (2026). Functionally Equivalent or Not? Graph-Grounded Differential Surrogate Execution for Code Equivalence. https://omanscience.com/ar/articles/functionally-equivalent-or-not-graph-grounded-differential-surrogate-execution-for-code-equivalence

MLA 9

Kachroo, Amit, et al. "Functionally Equivalent or Not? Graph-Grounded Differential Surrogate Execution for Code Equivalence." https://omanscience.com/ar/articles/functionally-equivalent-or-not-graph-grounded-differential-surrogate-execution-for-code-equivalence.

شيكاغو (المؤلف–التاريخ)

Kachroo, Amit, Like Hui, Haitao Mao, Yuhao Zhang, and Nguyen Vo. 2026. "Functionally Equivalent or Not? Graph-Grounded Differential Surrogate Execution for Code Equivalence." https://omanscience.com/ar/articles/functionally-equivalent-or-not-graph-grounded-differential-surrogate-execution-for-code-equivalence.

هارفارد

Kachroo, A., Hui, L., Mao, H., Zhang, Y. and Vo, N. (2026) 'Functionally Equivalent or Not? Graph-Grounded Differential Surrogate Execution for Code Equivalence', Available at: https://omanscience.com/ar/articles/functionally-equivalent-or-not-graph-grounded-differential-surrogate-execution-for-code-equivalence.

فانكوفر

Kachroo A, Hui L, Mao H, Zhang Y, Vo N. Functionally Equivalent or Not? Graph-Grounded Differential Surrogate Execution for Code Equivalence. https://omanscience.com/ar/articles/functionally-equivalent-or-not-graph-grounded-differential-surrogate-execution-for-code-equivalence

IEEE

A. Kachroo, L. Hui, H. Mao, Y. Zhang, and N. Vo, "Functionally Equivalent or Not? Graph-Grounded Differential Surrogate Execution for Code Equivalence," https://omanscience.com/ar/articles/functionally-equivalent-or-not-graph-grounded-differential-surrogate-execution-for-code-equivalence.