Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.compilers > #3742

Paper: An Extensive Empirical Study on Code Translation Technique

Path csiph.com!weretis.net!feeder9.news.weretis.net!news.misty.com!news.iecc.com!.POSTED.news.iecc.com!nerds-end
From John R Levine <johnl@taugh.com>
Newsgroups comp.compilers
Subject Paper: An Extensive Empirical Study on Code Translation Technique
Date Mon, 24 Aug 2026 14:46:47 -0400
Organization Compilers Central
Sender johnl%iecc.com
Approved comp.compilers@iecc.com
Message-ID <26-08-005@comp.compilers> (permalink)
MIME-Version 1.0
Content-Type text/plain; charset="UTF-8"
Injection-Info gal.iecc.com; posting-host="news.iecc.com:2001:470:1f07:1126:0:676f:7373:6970"; logging-data="74088"; mail-complaints-to="abuse@iecc.com"
Keywords paper, translator
Posted-Date 24 Aug 2026 14:47:10 EDT
X-submission-address compilers@iecc.com
X-moderator-address compilers-request@iecc.com
X-FAQ-and-archives http://compilers.iecc.com
Xref csiph.com comp.compilers:3742

Show key headers only | View raw


LLMs do a pretty good job of translating programming languages, but things
you would expect to be hard, like translating from dynamically to
statically typed languages, are indeed hard.

Abstract
Automated code translation is increasingly important for software
evolution, yet the relative strengths and limitations of learning-based
and large language model (LLM)-based techniques remain insufficiently
understood. To address this gap, we conduct a large-scale empirical study
comparing representative code translation techniques across methodological
paradigms and translation granularities. We evaluate learning-based
methods, LLM-based methods, and general-purpose LLMs on multilingual
method-level and class-level benchmarks involving multiple programming
languages. Our analysis considers executable correctness, code similarity,
translation direction, translation granularity, and failure patterns. The
results show that LLMs and LLM-based methods generally outperform
learning-based methods in method-level correctness, although similarity
metrics alone do not reliably reflect functional correctness. Translation
direction substantially affects performance, particularly when translating
between languages with different type-system characteristics. Class-level
translation remains considerably more difficult than method-level
translation because it requires preserving global semantics, interfaces,
member relationships, and cross-method dependencies. Our error analysis
further shows that static semantic errors and logical errors are the
primary challenges in existing code translation systems. These findings
provide empirical evidence and practical guidance for developing more
robust, type-aware, structure-aware, and context-aware code translation
techniques.

https://arxiv.org/abs/2608.20776

Regards,
John Levine, johnl@taugh.com, Taughannock Networks, Trumansburg NY
Please consider the environment before reading this e-mail. https://jl.ly

Back to comp.compilers | Previous | Next | Find similar | Unroll thread


Thread

Paper: An Extensive Empirical Study on Code Translation Technique John R Levine <johnl@taugh.com> - 2026-08-24 14:46 -0400

csiph-web