Skip to content

Repository files navigation

NanoEvolve

Automated neural architecture search using a multi-agent LLM system. Agents collaboratively ideate, implement, review, and evaluate transformer architecture variants — running the full loop autonomously on GPUs.

Built on top of nanoGPT for fast, reproducible training.

Key Results

Over 88 iterations of autonomous search, the system improved the baseline LLaMA-RoPE architecture:

Metric Baseline Best (exp_071) Improvement
Val Loss 3.4196 3.3492 -2.06%
Parameters ~124M ~81M -35%

Architecture Search Progress

How It Works

Each iteration runs a fully autonomous loop:

  1. Ideation — A moderator agent defines a research direction based on accumulated knowledge. Four domain-expert agents (physics, neuroscience, math, systems) propose modifications in parallel, then the moderator synthesizes them into a single concrete idea.
  2. Implementation — An implementation agent writes the PyTorch model code. A review agent checks for shape mismatches and correctness bugs. If the review fails, the implementation agent revises and resubmits (up to 3 retries).
  3. Training — A smoke test verifies the model compiles and runs without errors. If it passes, the model trains on FineWeb-Edu on a single GPU.
  4. Post-Training — An analysis agent compares the new result against the baseline and all previous experiments. A knowledge synthesizer updates a structured knowledge document with insights, which feeds back into the next iteration's ideation phase.

Training Setup

Parameter Value
Base architecture LLaMA-RoPE (12 layers, 12 heads, 768 dim)
Parameters ~124M (baseline)
Dataset FineWeb-Edu
Sequence length 512
Batch size 64
Gradient accumulation 2 steps
Max iterations 10,000 per experiment
Optimizer AdamW (lr=6e-4, cosine decay)
Precision bfloat16
GPUs 2x RTX 3090 (1 GPU per experiment, 2 experiments in parallel)

Discovered Techniques

The 6 experiments below drove the major val_loss improvements visible in the plot. All share a common design principle: identity-init — every new component starts as a no-op, preserving exact baseline behavior at initialization, and gradually activates during training.


1. Non-Uniform MLP Width Across Layers (exp_005: 3.4109)

Allocates MLP intermediate dimensions using a smooth monotonic ramp — narrowing early layers and widening later layers — while keeping total parameter count constant.

# In LlamaConfig.__post_init__:
alpha = self.mlp_alpha  # 0.7
n = self.n_layer
base = self.intermediate_size

raw_widths = [base * (alpha + (1.0 - alpha) * (2.0 * i / (n - 1))) for i in range(n)]

# Rescale so sum equals n * base (preserves total parameter count)
scale = (n * base) / sum(raw_widths)
self.intermediate_sizes = [128 * round(w * scale / 128) for w in raw_widths]

# In LlamaBlock.__init__:
self.mlp = LlamaMLP(config, config.intermediate_sizes[layer_idx])

+0 net params. Zero-cost reshuffling — early layers get narrow MLPs (pattern detection), later layers get wide MLPs (complex reasoning).


2. Learnable RMSNorm Temperature (exp_009: 3.4023)

Adds a learnable scalar temperature per RMSNorm layer, warm-started with a depth-dependent schedule. Early layers get softer normalization (tau > 1), deep layers get standard normalization (tau ≈ 1).

class DepthRMSNorm(nn.Module):
    def __init__(self, dim, layer_idx=0, n_layer=1, beta=0.25, eps=1e-6):
        super().__init__()
        self.eps = eps
        self.weight = nn.Parameter(torch.ones(dim))
        # Depth-dependent init: tau = 1 + beta*(1 - layer_idx/(n_layer-1))
        tau_init = 1.0 + beta * (1.0 - layer_idx / max(n_layer - 1, 1))
        self.tau = nn.Parameter(torch.tensor([tau_init]))

    def forward(self, x):
        norm = torch.rsqrt(x.pow(2).mean(-1, keepdim=True) + self.eps)
        return x * norm * self.weight * self.tau

+24 params (+0.00002%). One scalar per norm layer — the lightest possible modification.


3. Sigmoid-Gated Normalization Strength (exp_022: 3.3939)

Learns a sigmoid-gated scalar per RMSNorm that interpolates between identity (skip normalization) and full normalization, with a depth-dependent prior.

class RMSNorm(nn.Module):
    def __init__(self, dim, eps=1e-6, g_init=None):
        super().__init__()
        self.eps = eps
        self.weight = nn.Parameter(torch.ones(dim))
        self.g = nn.Parameter(torch.tensor(float(g_init))) if g_init is not None else None

    def forward(self, x):
        norm_x = x * torch.rsqrt(x.pow(2).mean(-1, keepdim=True) + self.eps)
        if self.g is not None:
            alpha = torch.sigmoid(self.g)
            x = x + alpha * (norm_x - x)  # lerp(identity, norm, alpha)
            return x * self.weight
        return norm_x * self.weight

# In LlamaBlock.__init__:
# g_init = 3.0 - 2.5 * (layer_idx / (n_layer - 1))
# Early: sigmoid(3.0) ≈ 0.95 (strong norm), Deep: sigmoid(0.5) ≈ 0.62 (partial norm)

+24 params (+0.00002%). Broke through the 3.40 barrier by letting each layer learn whether to normalize fully, partially, or nearly skip.


4. Post-Embedding Spectral Preconditioner (exp_044: 3.3876)

Inserts a learned low-rank (I + UV^T) transform right after embedding lookup, decoupling the input spectrum from the output-tied embedding matrix.

# In Llama.__init__:
r = config.embed_precond_rank  # 32
self.embed_U = nn.Parameter(torch.empty(config.n_embd, r))
self.embed_V = nn.Parameter(torch.zeros(r, config.n_embd))  # zero-init → starts as identity
nn.init.normal_(self.embed_U, mean=0.0, std=0.02 / math.sqrt(r))

# In Llama.forward:
x = self.embed_tokens(idx)
x = x + (x @ self.embed_U) @ self.embed_V  # (I + UV^T) spectral rotation

+49K params (+0.04%). With weight-tied embeddings, input and output share the same basis — this low-rank correction lets the transformer stack work in a rotated basis without breaking weight tying.


5. Learned Logit Temperature (exp_050: 3.3527)

Learns a per-token scalar temperature from the final hidden state to adaptively scale logits before softmax. The biggest single-step improvement (-0.035).

# In Llama.__init__:
# Zero-init → tau = softplus(0) + 0.5 ≈ 1.19 (near-identity)
self.w_temp = nn.Parameter(torch.zeros(config.n_embd))
self.b_temp = nn.Parameter(torch.zeros(1))

# In logit computation:
tau = F.softplus(
    torch.einsum('...d,d->...', hidden_states, self.w_temp) + self.b_temp
) + 0.5
logits = F.linear(hidden_states / tau.unsqueeze(-1), self.embed_tokens.weight)

+769 params (+0.0006%). High-entropy tokens (function words, punctuation) get sharper logits; low-entropy tokens (rare nouns, names) get softer distributions.


6. Attention Output Spectral Filter (exp_071: 3.3492)

Applies a zero-initialized low-rank (I + UV^T) correction to each layer's attention output, filtering noisy spectral components before the MLP sees them. The overall best.

# In LlamaBlock.__init__:
# Both zero-initialized → starts as identity (no-op)
self.attn_filter_U = nn.Parameter(torch.zeros(config.n_embd, 48))  # rank=48
self.attn_filter_V = nn.Parameter(torch.zeros(48, config.n_embd))

# In LlamaBlock.forward:
attn_out = self.attn(self.input_layernorm(x), cos, sin)
attn_out = attn_out + (attn_out @ self.attn_filter_U) @ self.attn_filter_V
h = x + self.attn_gate * attn_out

+884K params (+0.7%). Attention outputs contain noisy low-energy spectral directions that degrade MLP conditioning — this filter learns to suppress them.


Key Insights from 88 Experiments

  1. Low-rank (I + UV^T) corrections are the most effective pattern. The top results (exp_071, exp_044, exp_050) all apply learned identity perturbations at different points in the forward pass.

  2. Identity-init is essential. Every successful modification starts as a no-op. Depth-dependent or non-unit initialization consistently degrades performance (exp_002, exp_011, exp_018).

  3. Normalization is undertrained in standard transformers. Two of the six breakthroughs (exp_009, exp_022) simply added learnable scalars to RMSNorm — suggesting fixed normalization leaves performance on the table.

  4. Sequential attention→MLP is essential. Parallel sublayer execution consistently underperforms at this scale (exp_039, exp_040).

  5. torch.compile constrains architecture. Manual attention → OOM (exp_049). Dynamic shapes → recompilation OOM (exp_036). Matrix inverse → dispatch failure (exp_057).

  6. Low-rank factorization of Q/K/V hurts. Both Q/K (exp_027: 3.54) and V (exp_031: 3.44) low-rank factorizations degrade performance — full-rank projections are genuinely needed.


Installation

git clone https://github.com/smallhours19/nanoevolve.git
cd nanoevolve
pip install -r requirements.txt
python data/fineweb_edu/prepare.py

Prerequisites

  • Python 3.10+
  • CUDA-capable GPU(s) with ~24GB VRAM
  • Claude CLI installed and authenticated

Usage

# Run with default settings (5 iterations, 8 GPUs)
python orchestrator.py

# Customize the run
python orchestrator.py \
  --num-iterations 20 \
  --training-max-iters 10000 \
  --num-gpus 4 \
  --ideation-model opus \
  --impl-model opus

# Use specific GPUs
python orchestrator.py --gpu-ids 0,1,4,5

# Generate results plot
python plot_results.py

CLI Arguments

Argument Default Description
--num-iterations 5 Number of ideate-train cycles
--training-max-iters 10000 Max training steps per experiment
--num-gpus 8 Number of GPUs for parallel experiments
--gpu-ids None Specific GPU IDs (e.g., 0,1,4,5)
--ideation-model opus Claude model for ideation
--impl-model opus Claude model for implementation
--rate-limit-delay 5.0 Seconds between API calls

Acknowledgments

  • Training infrastructure based on nanoGPT by Andrej Karpathy
  • Architecture search powered by Claude via Claude CLI

About

No description, website, or topics provided.

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages