Research shelf / AI & machine learning / Neural Decompiler
Written Updated
Neural Decompiler — decompilation as conditional sequence modelling
Classical decompilation is a pipeline: disassembly, control-flow recovery, type reconstruction, pretty-printing. This work asks what happens if you treat the whole thing as one conditional sequence model — and is explicit that the answer is "a coherent trainable architecture", not a state-of-the-art recovery system.
Code exists and runs. Performance not independently checked.
An encoder–decoder Transformer with hierarchical memory and a load-balanced mixture of experts, reframing assembly-to-source recovery as a sequence modelling problem.
A PyTorch research implementation that reframes binary decompilation as conditional sequence modelling. Rather than a staged compiler pipeline run in reverse, an encoder–decoder Transformer maps assembly to source directly.
Three learnable mechanisms carry the design: a hierarchical memory module of learnable slots accessed by multi-head attention with gated fusion across levels; pre-norm GELU Transformer encoder and decoder layers with learned position embeddings; and a mixture of experts split into binary-focused and language-focused families.
The MoE routing is load-balanced by an auxiliary loss λ‖p̄ − u‖², added to a token cross-entropy with label smoothing. The bundled synthetic corpus validates the pipeline mechanically — forward pass, shape checks, loss decrease — and nothing more.
Every number, and what stands behind it
A claim is only worth the evidence attached to it. Each row below carries its basis: measured on the author’s own hardware, derived from the construction, measured on synthetic data, projected from literature, or simply cited.
| Claim | Figure | Basis | Context |
|---|---|---|---|
| Architecture | Encoder–decoder Transformer, pre-norm GELU | Derived | Implemented and runs |
| Hierarchical memory | Learnable slots, multi-head access, gated fusion | Derived | Implemented |
| Mixture of experts | Binary-focused and language-focused families | Derived | Implemented |
| MoE load balancing | Auxiliary loss λ‖p̄ − u‖² | Derived | Implemented |
| Pipeline validation | Forward pass, shape checks, loss decrease | Synthetic | On the bundled synthetic corpus only |
| Recovery quality on real binaries | Not claimed | Projected | Explicitly out of scope at this stage |
Measured — author-run experiment on the stated setup. Synthetic — measured, but on synthetic rather than real data. Derived — follows from the stated construction or proof. Projected — paper-stated projection, not an author-run benchmark. Cited — taken from external literature.
How it works
- Conditional sequence modelling. Assembly to source as one learned mapping, not a staged inversion.
- Gated hierarchical memory. Learnable slots read by attention, fused across levels by a gate.
- Load-balanced MoE. Routing regularised so experts are actually used.
- Synthetic corpus harness. Validates that the training loop is mechanically correct.
What it does not do
Taken from the folder’s own README. Nothing here has been softened.
- This is a trainable architecture and reference training loop, not a state-of-the-art recovery system on full binaries.
- Real deployment needs lifted-assembly plus source datasets with recorded compiler flags — which the project does not have.
- Long sequences and large vocabularies are not properly handled yet.
- Neural output is not type-checked or tested; nothing validates that the emitted source compiles or behaves.
- Dense mixtures may need replacing with sparse MoE kernels to be practical.
Free under AGPL-3.0+ for almost everyone
Personal use, charities, education and organisations under AUD 50,000 a year pay nothing. A tiered commercial licence covers everyone else.