Research shelf / AI & machine learning / Neural Decompiler

AI & machine learning

Neural Decompiler — decompilation as conditional sequence modelling

Classical decompilation is a pipeline: disassembly, control-flow recovery, type reconstruction, pretty-printing. This work asks what happens if you treat the whole thing as one conditional sequence model — and is explicit that the answer is "a coherent trainable architecture", not a state-of-the-art recovery system.

Reference implementation AGPL-3.0+ / commercial
Evidence level

Code exists and runs. Performance not independently checked.

FolderNeural Decompiler
FieldAI & machine learning
StatusReference implementation with training loop, inference harness and synthetic dataset validation.
What it is

An encoder–decoder Transformer with hierarchical memory and a load-balanced mixture of experts, reframing assembly-to-source recovery as a sequence modelling problem.

A PyTorch research implementation that reframes binary decompilation as conditional sequence modelling. Rather than a staged compiler pipeline run in reverse, an encoder–decoder Transformer maps assembly to source directly.

Three learnable mechanisms carry the design: a hierarchical memory module of learnable slots accessed by multi-head attention with gated fusion across levels; pre-norm GELU Transformer encoder and decoder layers with learned position embeddings; and a mixture of experts split into binary-focused and language-focused families.

The MoE routing is load-balanced by an auxiliary loss λ‖p̄ − u‖², added to a token cross-entropy with label smoothing. The bundled synthetic corpus validates the pipeline mechanically — forward pass, shape checks, loss decrease — and nothing more.

Compare with the shipping decompiler. RetDec Imortek takes the opposite bet: a deterministic decompiler is the auditable baseline, and neural refinement is optional, gated by a compile check, and off by default. This folder is the research line; that product is what survived contact with the requirement to be trustworthy.
Claims ledger

Every number, and what stands behind it

A claim is only worth the evidence attached to it. Each row below carries its basis: measured on the author’s own hardware, derived from the construction, measured on synthetic data, projected from literature, or simply cited.

Breakdown of this page’s claims by what stands behind each one
scroll to see the whole chart →
Every claim, weighted by its evidence. The table below is the same data row by row.
ClaimFigureBasisContext
ArchitectureEncoder–decoder Transformer, pre-norm GELUDerivedImplemented and runs
Hierarchical memoryLearnable slots, multi-head access, gated fusionDerivedImplemented
Mixture of expertsBinary-focused and language-focused familiesDerivedImplemented
MoE load balancingAuxiliary loss λ‖p̄ − u‖²DerivedImplemented
Pipeline validationForward pass, shape checks, loss decreaseSyntheticOn the bundled synthetic corpus only
Recovery quality on real binariesNot claimedProjectedExplicitly out of scope at this stage

Measured — author-run experiment on the stated setup. Synthetic — measured, but on synthetic rather than real data. Derived — follows from the stated construction or proof. Projected — paper-stated projection, not an author-run benchmark. Cited — taken from external literature.

Methods

How it works

  • Conditional sequence modelling. Assembly to source as one learned mapping, not a staged inversion.
  • Gated hierarchical memory. Learnable slots read by attention, fused across levels by a gate.
  • Load-balanced MoE. Routing regularised so experts are actually used.
  • Synthetic corpus harness. Validates that the training loop is mechanically correct.
Stated limitations

What it does not do

Taken from the folder’s own README. Nothing here has been softened.

  • This is a trainable architecture and reference training loop, not a state-of-the-art recovery system on full binaries.
  • Real deployment needs lifted-assembly plus source datasets with recorded compiler flags — which the project does not have.
  • Long sequences and large vocabularies are not properly handled yet.
  • Neural output is not type-checked or tested; nothing validates that the emitted source compiles or behaves.
  • Dense mixtures may need replacing with sparse MoE kernels to be practical.
Use it

Free under AGPL-3.0+ for almost everyone

Personal use, charities, education and organisations under AUD 50,000 a year pay nothing. A tiered commercial licence covers everyone else.