Skip to content
Project-LithonPublic

About

Statically-typed, Python-syntax language compiling directly to dependency-free x86-64 native code via dual JIT/AOT engines and an interpreter oracle. Engineered in C, C++, and x64 Assembly for C-class performance with Python readability.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

129 Commits

Folders and files

Repository files navigation

Lithon Logo

Lithon

Native execution for typed Python.

Build Status C++ Standard Architecture Typing License

Lithon is an experimental Python library and native x86-64 JIT that brings extended static typing, native execution, and explicit low-level memory access to Python code.


Lithon docs landing page


🚀 What is Lithon?

Lithon is a Python library and native execution engine that runs .py source code using an extended PEP-526-style static typing syntax.

Lithon is designed to keep Python's familiar source-code model while giving the compiler explicit type information that can be used for static verification, native machine-code generation, and predictable execution.

Lithon is not a separate programming language. Python remains the source language and .py remains the source format. Lithon extends the typing information available to the compiler so that supported Python code can be lowered to native x86-64 machine code.

For example:

x: int[64] = 10
y: int[64] = 20

result: int[64] = x + y

print(result)

Current status

Lithon is an experimental/community-preview project.

The compiler is already executing substantial typed programs natively and has a verification suite covering machine-code encoding, ABI correctness, static analysis, SSA, register allocation, floating point, containers, pointers, differential testing, and randomized fuzzing.

It is not yet a stable production library and should not be treated as a drop-in Python replacement.


✨ Why Lithon?

Lithon is built around several principles:

  • Python source with extended PEP-526 typing — Lithon works directly with .py source while providing additional type information to the native execution pipeline.
  • Mandatory static typing — types are verified before execution.
  • Fixed-width values — integer widths are explicit.
  • Native execution — supported programs can execute as generated x86-64 machine code.
  • No mandatory LLVM/Cranelift dependency — the backend contains its own machine-code encoder.
  • Interpreter fallback — unsupported native cases can still execute through Tier-0 when using automatic mode.
  • Explicit unsafe memory access — pointer variables require an _ prefix.
  • JIT/Compiler refusal over silent miscompilation — when Lithon cannot prove something, the native tier refuses it.

The last point is fundamental to the project:

If Lithon cannot prove that a program satisfies the rules required by its native backend, it should refuse native compilation rather than guess.


🧬 A Small Lithon Program

Static integers

a:int[64] = 20
b:int[64] = 22

result:int[64] = a + b

print(result)

Output:

42

The width is part of the type:

x:int[8] = 127
y:int[16] = 32000
z:int[64] = 9223372036854775807

Lithon does not silently turn these into arbitrary-precision Python integers.


🧠 Static Typing

Lithon's type checker is unconditional.

Removing annotations does not disable verification.

For example:

x = 10

is rejected when a declaration requires an explicit type.

Instead:

x:int[64] = 10

The compiler also checks:

  • definite assignment
  • branch type compatibility
  • return types
  • integer widths
  • narrowing conversions
  • function parameters
  • container element types
  • pointer types
  • shift bounds
  • container capacities

The principle is:

Unknown / unprovable
        │
        ▼
     REFUSE

rather than:

Unknown
  │
  ▼
Guess
  │
  ▼
Generate potentially incorrect machine code

🧷 Typed Pointers and Explicit Unsafe Variables

Lithon provides typed pointer values:

i:int[16] = 16
_ptrI:ptr[int[16]] = addressof(i)

print(_ptrI)
print(valueof(_ptrI))

Example output:

0x7fff330710e8
16

Why the _ matters

The _ prefix is mandatory for pointer variables.

_ptrI:ptr[int[16]] = addressof(i)

is valid.

A pointer variable without the required _ prefix is rejected and does not execute.

Likewise, a _-prefixed variable must satisfy the pointer rules.

This is Lithon's explicit unsafe-memory marker: the programmer can immediately see which variables belong to the raw/native memory domain.

The mechanism is intentionally different from Rust's unsafe blocks, but the design goal is similar:

Make potentially unsafe memory operations explicit instead of invisible.

Current pointer support includes:

  • ptr[T]
  • addressof()
  • valueof()
  • typed pointer arithmetic
  • compiler-enforced _ naming
  • static pointee typing
  • protection against pointer promotion into ordinary registers
  • frame-backed address stability

Pointers currently have deliberate restrictions. They cannot arbitrarily cross function boundaries, point into unsupported containers, or be used as a generic escape mechanism.


⚡ Native x86-64 Execution

Lithon's Tier-1 backend generates native x86-64 machine code directly.

The backend includes:

  • hand-written x86-64 instruction encoding
  • register allocation
  • liveness analysis
  • SSA / Phi handling
  • branch emission
  • stack-frame management
  • floating-point SSE2 operations
  • ABI-aware function calls
  • CPU feature detection
  • executable memory management

The generated code is placed into executable memory using the host operating system's native facilities such as mmap / VirtualAlloc.

Lithon does not require LLVM or Cranelift to generate its native machine code.


🏗️ Architecture

Lithon currently spans four implementation layers:

Layer Language Responsibility
Frontend Python Source syntax, annotations, AST → IR
Engine C++20 IR, type checking, analysis, optimization, execution
Bridge C ABI / native interoperability
Backend x86-64 Machine-code encoding and emission
graph TD
    A[Python .py Source] --> B[Frontend]
    B --> C[Typed IR]
    C --> D[Static Flow Verification]

    D -->|Native-safe| E[SSA / Optimization]
    E --> F[Register Allocation]
    F --> G[Hand-written x86-64 Encoder]
    G --> H[Executable Memory]
    H --> I[Native CPU Execution]

    D -->|Unsupported / unproven| J[Tier-0 C++ Interpreter]
Loading

🔀 Dual-Tier Execution

Lithon has two execution tiers.

Tier-1 — Native

The preferred path.

Lithon source
    ↓
Typed IR
    ↓
Static verification
    ↓
Optimization
    ↓
Register allocation
    ↓
x86-64 machine code
    ↓
CPU

Tier-0 — Interpreter

The fallback path.

If a program uses functionality that the native backend cannot currently support, automatic mode can execute it through the C++ interpreter.

This is useful during language development because unsupported features do not have to block the entire execution pipeline.

Strict mode

Use:

lithon program.py --strict

--strict means:

Native execution only. Refuse if Lithon cannot compile the program to the native tier.

This is particularly useful for testing the compiler because successful output cannot hide an interpreter fallback.


📦 Python Package

Lithon is being developed as a Python library / package, with a command-line interface for compiling and running Lithon programs.

The project is not yet released as a stable PyPI package.

During development, the repository can be installed in editable mode:

pip install -e .

Then:

lithon program.py

Useful commands:

lithon program.py
lithon program.py --strict
lithon program.py --ir
lithon program.py -v

The current development package layout is still being consolidated before the first public package release.


⚡ Quickstart

Requirements

  • CMake ≥ 3.20
  • C++20 compiler
  • Python ≥ 3.10
  • x86-64 Linux or Windows environment for the current native backend

The C++ engine itself does not depend on LLVM or Cranelift.

Build

git clone git@github.com:Project-Lithon/lithon.git
cd lithon

cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)

Run through the development CLI

pip install -e .

lithon tests/programs/float.py

Native-only:

lithon tests/programs/float.py --strict

Print IR:

lithon tests/programs/float.py --ir

Verbose execution information:

lithon tests/programs/float.py -v

🧪 Verification

Lithon places unusually strong emphasis on differential testing and machine-level verification.

The project does not consider "the program printed the expected number" enough to prove that the native backend is correct.

The verification suite checks several independent layers.

Current verification highlights

Area Current result
CTest 31 / 31 passed
x86-64 encoder comparisons 8,951 cases
Incorrect encoder cases 0
Typed differential fuzzing 300 / 300 passed
Bitwise / shift fuzzing 200 / 200 passed
Conditional-expression / SSA fuzzing 200 / 200 per suite
Direct float Phi sweep 38 matched / 0 mismatched
ABI modules audited 23
ABI stack/callee-saved failures 0
Typed regression suite 15 / 15 passed
General regression suite 23 / 23 passed

The full verification gate concludes:

ALL CHECKS PASSED

Encoder verification

The hand-written x86-64 encoder is compared against GNU as.

Current comparison:

8951 instruction cases
7661 byte-identical
1290 different but valid encodings
0 incorrect encodings

A byte difference is not automatically considered an error because x86-64 often has multiple valid encodings for the same instruction.

The important number is:

WRONG = 0

Differential testing

Native output is compared against the Tier-0 interpreter.

This catches errors where:

compiler accepts program
        ↓
machine code executes
        ↓
but machine code produces wrong result

Randomized programs are generated specifically to exercise:

  • arithmetic
  • control flow
  • SSA merges
  • loops
  • floating point
  • bitwise operations
  • shifts
  • modulo
  • register allocation
  • optimization passes

ABI verification

The native backend checks:

  • stack alignment
  • call boundaries
  • return boundaries
  • callee-saved register preservation

Compiled modules are disassembled and audited rather than merely assumed to follow the ABI.


🧮 Floating Point

Lithon currently supports float[64] through the native pipeline.

Supported operations include:

  • addition
  • subtraction
  • multiplication
  • division
  • modulo
  • comparisons
  • function arguments
  • function return values
  • native printing

The implementation uses SSE2 instructions and performs explicit handling for cases such as:

  • NaN comparisons
  • NaN divisors
  • positive and negative zero
  • division by zero
  • floating-point formatting
  • values crossing function calls

Float formatting is shared between the interpreter and native runtime so that both tiers produce identical textual output.


📚 Containers

Lists

Fixed-capacity lists are supported:

xs:list[int[32],4]

The capacity is static and the layout is packed according to element width.

Currently supported list operations include:

xs:list[int[32],4]

xs[0] = 10
xs[1] = 20

print(xs[0])
print(len(xs))

Bounds are statically checked when possible and dynamically trapped when the index is only known at runtime.

Tuples

Homogeneous immutable tuples are supported:

t:tuple[int[64],4] = (1, 2, 3, 4)

print(t[2])

Tuple construction must initialize all elements.

Writes after construction are rejected by the type checker.

Dictionaries

Fixed-capacity dictionaries use open addressing:

d:dict[int[64], int[64], 4] = {
    1: 10,
    2: 20
}

The table has a compile-time capacity and does not require a dynamically growing heap structure.

Supported operations include:

  • literal construction
  • lookup
  • contains
  • fixed-capacity hashing
  • collision resolution
  • deterministic miss traps

Mutation after construction is intentionally not currently supported.


🧠 SSA and Optimization

Lithon's compiler is progressively moving from a memory-based IR toward SSA.

Current components include:

  • CFG construction
  • dominator analysis
  • natural-loop analysis
  • loop canonicalization
  • Phi placement
  • Mem2Reg
  • SSA copy propagation
  • dead-value elimination
  • liveness analysis
  • interference-graph register allocation
  • Phi register coalescing
  • direct Phi lowering
  • constant folding
  • dead-code elimination
  • loop strength reduction
  • accumulator unrolling
  • optional floating-point reassociation

Direct Phi support

Phi handling is now implemented as an opt-in native path:

lithon program.py --strict

with the compiler's direct-Phi development path available through the relevant backend flags.

The default memory-based lowering remains available because the project uses the two paths as a differential-testing surface.

Floating-point reassociation

Lithon deliberately keeps floating-point reassociation off by default.

The opt-in:

--ffast-math-equivalent

can shorten floating-point dependency chains, but it may change the numerical result.

For example:

interpreter                     2.0
native                          2.0
native --ffast-math-equivalent  1.0

This behavior is intentional and documented rather than hidden behind a generic "optimization" switch.


% Is Not Python's %

Lithon deliberately uses C-style truncating remainder semantics.

print(-7 % 3)   # Lithon: -1
print(7 % -3)   # Lithon: 1

CPython instead uses floor-division semantics:

Lithon / C-style:  remainder follows the dividend
Python:            remainder follows the divisor

This is an intentional language-design decision.


🔀 Integer Operations and Shifts

Integer operations operate on fixed-width machine integers.

Supported bitwise operators:

&
|
^
<<
>>

Right shift is arithmetic.

Shift counts must remain within the valid machine range:

0 <= shift < 64

A constant invalid shift is rejected at compile time, while a dynamic invalid shift traps at runtime rather than relying on x86's implicit masking behavior.

This keeps the language semantics explicit instead of accidentally inheriting the CPU's register-level behavior.


🛡️ Safety Philosophy

Lithon is not trying to eliminate low-level programming.

It is trying to make low-level programming explicit and statically constrained.

The compiler therefore prefers:

prove → compile

over:

guess → compile → hope

Examples include:

  • mandatory variable annotations
  • fixed-width integers
  • static narrowing checks
  • definite-assignment analysis
  • fixed container capacities
  • typed pointers
  • explicit _ pointer naming
  • compile-time shift validation
  • ABI verification
  • native/interpreter differential testing

The unsafe-memory model is intentionally visible:

_ptr:ptr[int[64]]

rather than hiding the fact that a variable contains a raw address.


🆚 How Is Lithon Different?

Feature CPython Cython / mypyc PyPy Lithon
Syntax Python Python Python Python .py + Extended PEP-526
Primary execution Bytecode interpreter C extensions Tracing JIT Native x86-64
Static typing No Optional No Mandatory
Fixed-width integers No Possible No Built-in
Raw typed pointers No Via C No Built-in
Native backend No C/C++ JIT runtime Custom x86-64 emitter
LLVM required No No No No
Interpreter fallback Yes N/A Yes Yes
SSA compiler pipeline No External/compiler-dependent Internal JIT Yes
Current target Cross-platform Cross-platform Cross-platform x86-64

Lithon is not intended to replace Python's enormous ecosystem.

Instead, it explores a different point in the design space:

What if Python code could retain its familiar source model while giving the execution engine explicit static types, fixed-width values, native execution, and controlled low-level memory access?


🎯 Potential Use Cases

Lithon is particularly interesting for programs where predictable native execution matters:

  • numeric algorithms
  • tight computational loops
  • low-latency utilities
  • systems-oriented tooling
  • performance-sensitive data processing
  • experimental compiler research
  • native execution experiments
  • educational work involving compilers and machine code

Performance claims should always be treated as workload-dependent.

Lithon is not currently claiming that every Python program will run faster than CPython, PyPy, Cython, Rust, or C++.


🧭 Roadmap

              Current
                 │
                 ▼
       ┌─────────────────────┐
       │ Dual-Tier Native JIT│
       │ Static Type System  │
       │ SSA / Optimizations │
       │ x86-64 Backend      │
       └──────────┬──────────┘
                  │
                  ▼
       ┌─────────────────────┐
       │ AOT Compilation     │
       │ Standalone Binaries │
       └──────────┬──────────┘
                  │
                  ▼
       ┌─────────────────────┐
       │ Hardware / Systems  │
       │ Integration         │
       │ SIMD / Syscalls     │
       └─────────────────────┘

Phase I — Native compiler

  • Dual-tier execution
  • Static type flow verification
  • Hand-written x86-64 encoder
  • Register allocation
  • SSA pipeline
  • Floating-point pipeline
  • Lists
  • Tuples
  • Fixed dictionaries
  • Typed pointers
  • ABI verification
  • Differential testing
  • Randomized compiler fuzzing
  • CPU feature detection

Phase II — AOT

  • ELF64 generation
  • PE32+ generation
  • Standalone native executables
  • Runtime-independent deployment

Phase III — Systems / hardware integration

  • Direct syscall emission
  • Expanded raw-memory operations
  • VEX instruction encoding
  • AVX2 / SIMD vectorization
  • SIMD reductions
  • Expanded native FFI
  • ARM64 backend

📍 Current Known Limitations

Lithon is powerful on its supported subset, but its native execution engine is still experimental.

Current limitations include:

  • x86-64 is the primary native backend.
  • ARM64 is not yet implemented.
  • More than two function arguments are not yet supported by the current native calling pipeline.
  • Some container operations remain intentionally restricted.
  • Pointer values have strict restrictions around function boundaries and container access.
  • Some SSA/native paths remain opt-in or have known unsupported cases.
  • There is no production AOT compiler yet.
  • The PyPI package has not yet received its stable public release.
  • Lithon is not intended to be a drop-in CPython replacement.

When the compiler cannot prove a construct is supported, it should refuse the native tier rather than silently produce questionable machine code.


🧪 Testing

Build

cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)

Unit tests

ctest --test-dir build --output-on-failure

Full verification

bash tools/verify_all.sh

Regression suites

python3 tools/run_regression.py
python3 tools/run_typed_regression.py

Tier differential testing

python3 tools/run_tier_diff.py

Differential fuzzing

python3 tools/fuzz_diff.py --count 300
python3 tools/fuzz_diff.py --count 300 --floats
python3 tools/fuzz_diff.py --count 300 --bitwise
python3 tools/fuzz_diff.py --count 300 --phi
python3 tools/fuzz_diff.py --mod --count 300

Benchmarking

python3 tools/native_bench.py --runs 30 --pin 2

For reliable comparisons, pin the benchmark to a CPU and compare results generated by the same benchmark harness.


📊 Latest Benchmark

⚡ Latest Benchmark

Benchmark Lithon JIT Reference Speedup Result
fib(30) 7.1495 ms 1563.1602 ms 218.6× 832040

Status: PASS

Last updated by Lithon Reporter Mamba.

Benchmark results are workload- and machine-dependent. They should be reproduced locally before being used as general performance claims.


🤝 Community

Lithon is entering its community-preview stage.

If you are interested in:

  • compiler construction
  • JIT compilation
  • x86-64 machine-code generation
  • static analysis
  • SSA and register allocation
  • systems programming

then feedback, experiments, bug reports, and technical criticism are welcome.

The project is still evolving. In particular, feedback on the type system, pointer/unsafe model, compiler architecture, and language semantics is especially valuable.


📜 License

Lithon is released under the MIT License.

Native execution for typed Python. Explicit by design.

About

Statically-typed, Python-syntax language compiling directly to dependency-free x86-64 native code via dual JIT/AOT engines and an interpreter oracle. Engineered in C, C++, and x64 Assembly for C-class performance with Python readability.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages