Lithon is an experimental Python library and native x86-64 JIT that brings extended static typing, native execution, and explicit low-level memory access to Python code.
Lithon is a Python library and native execution engine that runs .py
source code using an extended PEP-526-style static typing syntax.
Lithon is designed to keep Python's familiar source-code model while giving the compiler explicit type information that can be used for static verification, native machine-code generation, and predictable execution.
Lithon is not a separate programming language. Python remains the source
language and .py remains the source format. Lithon extends the typing
information available to the compiler so that supported Python code can be
lowered to native x86-64 machine code.
For example:
x: int[64] = 10
y: int[64] = 20
result: int[64] = x + y
print(result)Lithon is an experimental/community-preview project.
The compiler is already executing substantial typed programs natively and has a verification suite covering machine-code encoding, ABI correctness, static analysis, SSA, register allocation, floating point, containers, pointers, differential testing, and randomized fuzzing.
It is not yet a stable production library and should not be treated as a drop-in Python replacement.
Lithon is built around several principles:
- Python source with extended PEP-526 typing — Lithon works directly with
.pysource while providing additional type information to the native execution pipeline. - Mandatory static typing — types are verified before execution.
- Fixed-width values — integer widths are explicit.
- Native execution — supported programs can execute as generated x86-64 machine code.
- No mandatory LLVM/Cranelift dependency — the backend contains its own machine-code encoder.
- Interpreter fallback — unsupported native cases can still execute through Tier-0 when using automatic mode.
- Explicit unsafe memory access — pointer variables require an
_prefix. - JIT/Compiler refusal over silent miscompilation — when Lithon cannot prove something, the native tier refuses it.
The last point is fundamental to the project:
If Lithon cannot prove that a program satisfies the rules required by its native backend, it should refuse native compilation rather than guess.
a:int[64] = 20
b:int[64] = 22
result:int[64] = a + b
print(result)Output:
42
The width is part of the type:
x:int[8] = 127
y:int[16] = 32000
z:int[64] = 9223372036854775807Lithon does not silently turn these into arbitrary-precision Python integers.
Lithon's type checker is unconditional.
Removing annotations does not disable verification.
For example:
x = 10is rejected when a declaration requires an explicit type.
Instead:
x:int[64] = 10The compiler also checks:
- definite assignment
- branch type compatibility
- return types
- integer widths
- narrowing conversions
- function parameters
- container element types
- pointer types
- shift bounds
- container capacities
The principle is:
Unknown / unprovable
│
▼
REFUSE
rather than:
Unknown
│
▼
Guess
│
▼
Generate potentially incorrect machine code
Lithon provides typed pointer values:
i:int[16] = 16
_ptrI:ptr[int[16]] = addressof(i)
print(_ptrI)
print(valueof(_ptrI))Example output:
0x7fff330710e8
16
The _ prefix is mandatory for pointer variables.
_ptrI:ptr[int[16]] = addressof(i)is valid.
A pointer variable without the required _ prefix is rejected and does not
execute.
Likewise, a _-prefixed variable must satisfy the pointer rules.
This is Lithon's explicit unsafe-memory marker: the programmer can immediately see which variables belong to the raw/native memory domain.
The mechanism is intentionally different from Rust's unsafe blocks, but the
design goal is similar:
Make potentially unsafe memory operations explicit instead of invisible.
Current pointer support includes:
ptr[T]addressof()valueof()- typed pointer arithmetic
- compiler-enforced
_naming - static pointee typing
- protection against pointer promotion into ordinary registers
- frame-backed address stability
Pointers currently have deliberate restrictions. They cannot arbitrarily cross function boundaries, point into unsupported containers, or be used as a generic escape mechanism.
Lithon's Tier-1 backend generates native x86-64 machine code directly.
The backend includes:
- hand-written x86-64 instruction encoding
- register allocation
- liveness analysis
- SSA / Phi handling
- branch emission
- stack-frame management
- floating-point SSE2 operations
- ABI-aware function calls
- CPU feature detection
- executable memory management
The generated code is placed into executable memory using the host operating
system's native facilities such as mmap / VirtualAlloc.
Lithon does not require LLVM or Cranelift to generate its native machine code.
Lithon currently spans four implementation layers:
| Layer | Language | Responsibility |
|---|---|---|
| Frontend | Python | Source syntax, annotations, AST → IR |
| Engine | C++20 | IR, type checking, analysis, optimization, execution |
| Bridge | C | ABI / native interoperability |
| Backend | x86-64 | Machine-code encoding and emission |
graph TD
A[Python .py Source] --> B[Frontend]
B --> C[Typed IR]
C --> D[Static Flow Verification]
D -->|Native-safe| E[SSA / Optimization]
E --> F[Register Allocation]
F --> G[Hand-written x86-64 Encoder]
G --> H[Executable Memory]
H --> I[Native CPU Execution]
D -->|Unsupported / unproven| J[Tier-0 C++ Interpreter]
Lithon has two execution tiers.
The preferred path.
Lithon source
↓
Typed IR
↓
Static verification
↓
Optimization
↓
Register allocation
↓
x86-64 machine code
↓
CPU
The fallback path.
If a program uses functionality that the native backend cannot currently support, automatic mode can execute it through the C++ interpreter.
This is useful during language development because unsupported features do not have to block the entire execution pipeline.
Use:
lithon program.py --strict--strict means:
Native execution only. Refuse if Lithon cannot compile the program to the native tier.
This is particularly useful for testing the compiler because successful output cannot hide an interpreter fallback.
Lithon is being developed as a Python library / package, with a command-line interface for compiling and running Lithon programs.
The project is not yet released as a stable PyPI package.
During development, the repository can be installed in editable mode:
pip install -e .Then:
lithon program.pyUseful commands:
lithon program.py
lithon program.py --strict
lithon program.py --ir
lithon program.py -vThe current development package layout is still being consolidated before the first public package release.
- CMake ≥ 3.20
- C++20 compiler
- Python ≥ 3.10
- x86-64 Linux or Windows environment for the current native backend
The C++ engine itself does not depend on LLVM or Cranelift.
git clone git@github.com:Project-Lithon/lithon.git
cd lithon
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)pip install -e .
lithon tests/programs/float.pyNative-only:
lithon tests/programs/float.py --strictPrint IR:
lithon tests/programs/float.py --irVerbose execution information:
lithon tests/programs/float.py -vLithon places unusually strong emphasis on differential testing and machine-level verification.
The project does not consider "the program printed the expected number" enough to prove that the native backend is correct.
The verification suite checks several independent layers.
| Area | Current result |
|---|---|
| CTest | 31 / 31 passed |
| x86-64 encoder comparisons | 8,951 cases |
| Incorrect encoder cases | 0 |
| Typed differential fuzzing | 300 / 300 passed |
| Bitwise / shift fuzzing | 200 / 200 passed |
| Conditional-expression / SSA fuzzing | 200 / 200 per suite |
| Direct float Phi sweep | 38 matched / 0 mismatched |
| ABI modules audited | 23 |
| ABI stack/callee-saved failures | 0 |
| Typed regression suite | 15 / 15 passed |
| General regression suite | 23 / 23 passed |
The full verification gate concludes:
ALL CHECKS PASSED
The hand-written x86-64 encoder is compared against GNU as.
Current comparison:
8951 instruction cases
7661 byte-identical
1290 different but valid encodings
0 incorrect encodings
A byte difference is not automatically considered an error because x86-64 often has multiple valid encodings for the same instruction.
The important number is:
WRONG = 0
Native output is compared against the Tier-0 interpreter.
This catches errors where:
compiler accepts program
↓
machine code executes
↓
but machine code produces wrong result
Randomized programs are generated specifically to exercise:
- arithmetic
- control flow
- SSA merges
- loops
- floating point
- bitwise operations
- shifts
- modulo
- register allocation
- optimization passes
The native backend checks:
- stack alignment
- call boundaries
- return boundaries
- callee-saved register preservation
Compiled modules are disassembled and audited rather than merely assumed to follow the ABI.
Lithon currently supports float[64] through the native pipeline.
Supported operations include:
- addition
- subtraction
- multiplication
- division
- modulo
- comparisons
- function arguments
- function return values
- native printing
The implementation uses SSE2 instructions and performs explicit handling for cases such as:
- NaN comparisons
- NaN divisors
- positive and negative zero
- division by zero
- floating-point formatting
- values crossing function calls
Float formatting is shared between the interpreter and native runtime so that both tiers produce identical textual output.
Fixed-capacity lists are supported:
xs:list[int[32],4]The capacity is static and the layout is packed according to element width.
Currently supported list operations include:
xs:list[int[32],4]
xs[0] = 10
xs[1] = 20
print(xs[0])
print(len(xs))Bounds are statically checked when possible and dynamically trapped when the index is only known at runtime.
Homogeneous immutable tuples are supported:
t:tuple[int[64],4] = (1, 2, 3, 4)
print(t[2])Tuple construction must initialize all elements.
Writes after construction are rejected by the type checker.
Fixed-capacity dictionaries use open addressing:
d:dict[int[64], int[64], 4] = {
1: 10,
2: 20
}The table has a compile-time capacity and does not require a dynamically growing heap structure.
Supported operations include:
- literal construction
- lookup
contains- fixed-capacity hashing
- collision resolution
- deterministic miss traps
Mutation after construction is intentionally not currently supported.
Lithon's compiler is progressively moving from a memory-based IR toward SSA.
Current components include:
- CFG construction
- dominator analysis
- natural-loop analysis
- loop canonicalization
- Phi placement
- Mem2Reg
- SSA copy propagation
- dead-value elimination
- liveness analysis
- interference-graph register allocation
- Phi register coalescing
- direct Phi lowering
- constant folding
- dead-code elimination
- loop strength reduction
- accumulator unrolling
- optional floating-point reassociation
Phi handling is now implemented as an opt-in native path:
lithon program.py --strictwith the compiler's direct-Phi development path available through the relevant backend flags.
The default memory-based lowering remains available because the project uses the two paths as a differential-testing surface.
Lithon deliberately keeps floating-point reassociation off by default.
The opt-in:
--ffast-math-equivalentcan shorten floating-point dependency chains, but it may change the numerical result.
For example:
interpreter 2.0
native 2.0
native --ffast-math-equivalent 1.0
This behavior is intentional and documented rather than hidden behind a generic "optimization" switch.
Lithon deliberately uses C-style truncating remainder semantics.
print(-7 % 3) # Lithon: -1
print(7 % -3) # Lithon: 1CPython instead uses floor-division semantics:
Lithon / C-style: remainder follows the dividend
Python: remainder follows the divisor
This is an intentional language-design decision.
Integer operations operate on fixed-width machine integers.
Supported bitwise operators:
&
|
^
<<
>>Right shift is arithmetic.
Shift counts must remain within the valid machine range:
0 <= shift < 64
A constant invalid shift is rejected at compile time, while a dynamic invalid shift traps at runtime rather than relying on x86's implicit masking behavior.
This keeps the language semantics explicit instead of accidentally inheriting the CPU's register-level behavior.
Lithon is not trying to eliminate low-level programming.
It is trying to make low-level programming explicit and statically constrained.
The compiler therefore prefers:
prove → compile
over:
guess → compile → hope
Examples include:
- mandatory variable annotations
- fixed-width integers
- static narrowing checks
- definite-assignment analysis
- fixed container capacities
- typed pointers
- explicit
_pointer naming - compile-time shift validation
- ABI verification
- native/interpreter differential testing
The unsafe-memory model is intentionally visible:
_ptr:ptr[int[64]]rather than hiding the fact that a variable contains a raw address.
| Feature | CPython | Cython / mypyc | PyPy | Lithon |
|---|---|---|---|---|
| Syntax | Python | Python | Python | Python .py + Extended PEP-526 |
| Primary execution | Bytecode interpreter | C extensions | Tracing JIT | Native x86-64 |
| Static typing | No | Optional | No | Mandatory |
| Fixed-width integers | No | Possible | No | Built-in |
| Raw typed pointers | No | Via C | No | Built-in |
| Native backend | No | C/C++ | JIT runtime | Custom x86-64 emitter |
| LLVM required | No | No | No | No |
| Interpreter fallback | Yes | N/A | Yes | Yes |
| SSA compiler pipeline | No | External/compiler-dependent | Internal JIT | Yes |
| Current target | Cross-platform | Cross-platform | Cross-platform | x86-64 |
Lithon is not intended to replace Python's enormous ecosystem.
Instead, it explores a different point in the design space:
What if Python code could retain its familiar source model while giving the execution engine explicit static types, fixed-width values, native execution, and controlled low-level memory access?
Lithon is particularly interesting for programs where predictable native execution matters:
- numeric algorithms
- tight computational loops
- low-latency utilities
- systems-oriented tooling
- performance-sensitive data processing
- experimental compiler research
- native execution experiments
- educational work involving compilers and machine code
Performance claims should always be treated as workload-dependent.
Lithon is not currently claiming that every Python program will run faster than CPython, PyPy, Cython, Rust, or C++.
Current
│
▼
┌─────────────────────┐
│ Dual-Tier Native JIT│
│ Static Type System │
│ SSA / Optimizations │
│ x86-64 Backend │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ AOT Compilation │
│ Standalone Binaries │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ Hardware / Systems │
│ Integration │
│ SIMD / Syscalls │
└─────────────────────┘
- Dual-tier execution
- Static type flow verification
- Hand-written x86-64 encoder
- Register allocation
- SSA pipeline
- Floating-point pipeline
- Lists
- Tuples
- Fixed dictionaries
- Typed pointers
- ABI verification
- Differential testing
- Randomized compiler fuzzing
- CPU feature detection
- ELF64 generation
- PE32+ generation
- Standalone native executables
- Runtime-independent deployment
- Direct syscall emission
- Expanded raw-memory operations
- VEX instruction encoding
- AVX2 / SIMD vectorization
- SIMD reductions
- Expanded native FFI
- ARM64 backend
Lithon is powerful on its supported subset, but its native execution engine is still experimental.
Current limitations include:
- x86-64 is the primary native backend.
- ARM64 is not yet implemented.
- More than two function arguments are not yet supported by the current native calling pipeline.
- Some container operations remain intentionally restricted.
- Pointer values have strict restrictions around function boundaries and container access.
- Some SSA/native paths remain opt-in or have known unsupported cases.
- There is no production AOT compiler yet.
- The PyPI package has not yet received its stable public release.
- Lithon is not intended to be a drop-in CPython replacement.
When the compiler cannot prove a construct is supported, it should refuse the native tier rather than silently produce questionable machine code.
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)ctest --test-dir build --output-on-failurebash tools/verify_all.shpython3 tools/run_regression.py
python3 tools/run_typed_regression.pypython3 tools/run_tier_diff.pypython3 tools/fuzz_diff.py --count 300
python3 tools/fuzz_diff.py --count 300 --floats
python3 tools/fuzz_diff.py --count 300 --bitwise
python3 tools/fuzz_diff.py --count 300 --phi
python3 tools/fuzz_diff.py --mod --count 300python3 tools/native_bench.py --runs 30 --pin 2For reliable comparisons, pin the benchmark to a CPU and compare results generated by the same benchmark harness.
| Benchmark | Lithon JIT | Reference | Speedup | Result |
|---|---|---|---|---|
| fib(30) | 7.1495 ms | 1563.1602 ms | 218.6× | 832040 |
Status: PASS
Last updated by Lithon Reporter Mamba.
Benchmark results are workload- and machine-dependent. They should be reproduced locally before being used as general performance claims.
Lithon is entering its community-preview stage.
If you are interested in:
- compiler construction
- JIT compilation
- x86-64 machine-code generation
- static analysis
- SSA and register allocation
- systems programming
then feedback, experiments, bug reports, and technical criticism are welcome.
The project is still evolving. In particular, feedback on the type system, pointer/unsafe model, compiler architecture, and language semantics is especially valuable.
Lithon is released under the MIT License.
Native execution for typed Python. Explicit by design.
