Automated neural architecture search using a multi-agent LLM system. Agents collaboratively ideate, implement, review, and evaluate transformer architecture variants — running the full loop autonomously on GPUs.
Built on top of nanoGPT for fast, reproducible training.
Over 88 iterations of autonomous search, the system improved the baseline LLaMA-RoPE architecture:
| Metric | Baseline | Best (exp_071) | Improvement |
|---|---|---|---|
| Val Loss | 3.4196 | 3.3492 | -2.06% |
| Parameters | ~124M | ~81M | -35% |
Each iteration runs a fully autonomous loop:
- Ideation — A moderator agent defines a research direction based on accumulated knowledge. Four domain-expert agents (physics, neuroscience, math, systems) propose modifications in parallel, then the moderator synthesizes them into a single concrete idea.
- Implementation — An implementation agent writes the PyTorch model code. A review agent checks for shape mismatches and correctness bugs. If the review fails, the implementation agent revises and resubmits (up to 3 retries).
- Training — A smoke test verifies the model compiles and runs without errors. If it passes, the model trains on FineWeb-Edu on a single GPU.
- Post-Training — An analysis agent compares the new result against the baseline and all previous experiments. A knowledge synthesizer updates a structured knowledge document with insights, which feeds back into the next iteration's ideation phase.
| Parameter | Value |
|---|---|
| Base architecture | LLaMA-RoPE (12 layers, 12 heads, 768 dim) |
| Parameters | ~124M (baseline) |
| Dataset | FineWeb-Edu |
| Sequence length | 512 |
| Batch size | 64 |
| Gradient accumulation | 2 steps |
| Max iterations | 10,000 per experiment |
| Optimizer | AdamW (lr=6e-4, cosine decay) |
| Precision | bfloat16 |
| GPUs | 2x RTX 3090 (1 GPU per experiment, 2 experiments in parallel) |
The 6 experiments below drove the major val_loss improvements visible in the plot. All share a common design principle: identity-init — every new component starts as a no-op, preserving exact baseline behavior at initialization, and gradually activates during training.
Allocates MLP intermediate dimensions using a smooth monotonic ramp — narrowing early layers and widening later layers — while keeping total parameter count constant.
# In LlamaConfig.__post_init__:
alpha = self.mlp_alpha # 0.7
n = self.n_layer
base = self.intermediate_size
raw_widths = [base * (alpha + (1.0 - alpha) * (2.0 * i / (n - 1))) for i in range(n)]
# Rescale so sum equals n * base (preserves total parameter count)
scale = (n * base) / sum(raw_widths)
self.intermediate_sizes = [128 * round(w * scale / 128) for w in raw_widths]
# In LlamaBlock.__init__:
self.mlp = LlamaMLP(config, config.intermediate_sizes[layer_idx])+0 net params. Zero-cost reshuffling — early layers get narrow MLPs (pattern detection), later layers get wide MLPs (complex reasoning).
Adds a learnable scalar temperature per RMSNorm layer, warm-started with a depth-dependent schedule. Early layers get softer normalization (tau > 1), deep layers get standard normalization (tau ≈ 1).
class DepthRMSNorm(nn.Module):
def __init__(self, dim, layer_idx=0, n_layer=1, beta=0.25, eps=1e-6):
super().__init__()
self.eps = eps
self.weight = nn.Parameter(torch.ones(dim))
# Depth-dependent init: tau = 1 + beta*(1 - layer_idx/(n_layer-1))
tau_init = 1.0 + beta * (1.0 - layer_idx / max(n_layer - 1, 1))
self.tau = nn.Parameter(torch.tensor([tau_init]))
def forward(self, x):
norm = torch.rsqrt(x.pow(2).mean(-1, keepdim=True) + self.eps)
return x * norm * self.weight * self.tau+24 params (+0.00002%). One scalar per norm layer — the lightest possible modification.
Learns a sigmoid-gated scalar per RMSNorm that interpolates between identity (skip normalization) and full normalization, with a depth-dependent prior.
class RMSNorm(nn.Module):
def __init__(self, dim, eps=1e-6, g_init=None):
super().__init__()
self.eps = eps
self.weight = nn.Parameter(torch.ones(dim))
self.g = nn.Parameter(torch.tensor(float(g_init))) if g_init is not None else None
def forward(self, x):
norm_x = x * torch.rsqrt(x.pow(2).mean(-1, keepdim=True) + self.eps)
if self.g is not None:
alpha = torch.sigmoid(self.g)
x = x + alpha * (norm_x - x) # lerp(identity, norm, alpha)
return x * self.weight
return norm_x * self.weight
# In LlamaBlock.__init__:
# g_init = 3.0 - 2.5 * (layer_idx / (n_layer - 1))
# Early: sigmoid(3.0) ≈ 0.95 (strong norm), Deep: sigmoid(0.5) ≈ 0.62 (partial norm)+24 params (+0.00002%). Broke through the 3.40 barrier by letting each layer learn whether to normalize fully, partially, or nearly skip.
Inserts a learned low-rank (I + UV^T) transform right after embedding lookup, decoupling the input spectrum from the output-tied embedding matrix.
# In Llama.__init__:
r = config.embed_precond_rank # 32
self.embed_U = nn.Parameter(torch.empty(config.n_embd, r))
self.embed_V = nn.Parameter(torch.zeros(r, config.n_embd)) # zero-init → starts as identity
nn.init.normal_(self.embed_U, mean=0.0, std=0.02 / math.sqrt(r))
# In Llama.forward:
x = self.embed_tokens(idx)
x = x + (x @ self.embed_U) @ self.embed_V # (I + UV^T) spectral rotation+49K params (+0.04%). With weight-tied embeddings, input and output share the same basis — this low-rank correction lets the transformer stack work in a rotated basis without breaking weight tying.
Learns a per-token scalar temperature from the final hidden state to adaptively scale logits before softmax. The biggest single-step improvement (-0.035).
# In Llama.__init__:
# Zero-init → tau = softplus(0) + 0.5 ≈ 1.19 (near-identity)
self.w_temp = nn.Parameter(torch.zeros(config.n_embd))
self.b_temp = nn.Parameter(torch.zeros(1))
# In logit computation:
tau = F.softplus(
torch.einsum('...d,d->...', hidden_states, self.w_temp) + self.b_temp
) + 0.5
logits = F.linear(hidden_states / tau.unsqueeze(-1), self.embed_tokens.weight)+769 params (+0.0006%). High-entropy tokens (function words, punctuation) get sharper logits; low-entropy tokens (rare nouns, names) get softer distributions.
Applies a zero-initialized low-rank (I + UV^T) correction to each layer's attention output, filtering noisy spectral components before the MLP sees them. The overall best.
# In LlamaBlock.__init__:
# Both zero-initialized → starts as identity (no-op)
self.attn_filter_U = nn.Parameter(torch.zeros(config.n_embd, 48)) # rank=48
self.attn_filter_V = nn.Parameter(torch.zeros(48, config.n_embd))
# In LlamaBlock.forward:
attn_out = self.attn(self.input_layernorm(x), cos, sin)
attn_out = attn_out + (attn_out @ self.attn_filter_U) @ self.attn_filter_V
h = x + self.attn_gate * attn_out+884K params (+0.7%). Attention outputs contain noisy low-energy spectral directions that degrade MLP conditioning — this filter learns to suppress them.
-
Low-rank
(I + UV^T)corrections are the most effective pattern. The top results (exp_071, exp_044, exp_050) all apply learned identity perturbations at different points in the forward pass. -
Identity-init is essential. Every successful modification starts as a no-op. Depth-dependent or non-unit initialization consistently degrades performance (exp_002, exp_011, exp_018).
-
Normalization is undertrained in standard transformers. Two of the six breakthroughs (exp_009, exp_022) simply added learnable scalars to RMSNorm — suggesting fixed normalization leaves performance on the table.
-
Sequential attention→MLP is essential. Parallel sublayer execution consistently underperforms at this scale (exp_039, exp_040).
-
torch.compile constrains architecture. Manual attention → OOM (exp_049). Dynamic shapes → recompilation OOM (exp_036). Matrix inverse → dispatch failure (exp_057).
-
Low-rank factorization of Q/K/V hurts. Both Q/K (exp_027: 3.54) and V (exp_031: 3.44) low-rank factorizations degrade performance — full-rank projections are genuinely needed.
git clone https://github.com/smallhours19/nanoevolve.git
cd nanoevolve
pip install -r requirements.txt
python data/fineweb_edu/prepare.py- Python 3.10+
- CUDA-capable GPU(s) with ~24GB VRAM
- Claude CLI installed and authenticated
# Run with default settings (5 iterations, 8 GPUs)
python orchestrator.py
# Customize the run
python orchestrator.py \
--num-iterations 20 \
--training-max-iters 10000 \
--num-gpus 4 \
--ideation-model opus \
--impl-model opus
# Use specific GPUs
python orchestrator.py --gpu-ids 0,1,4,5
# Generate results plot
python plot_results.py| Argument | Default | Description |
|---|---|---|
--num-iterations |
5 | Number of ideate-train cycles |
--training-max-iters |
10000 | Max training steps per experiment |
--num-gpus |
8 | Number of GPUs for parallel experiments |
--gpu-ids |
None | Specific GPU IDs (e.g., 0,1,4,5) |
--ideation-model |
opus | Claude model for ideation |
--impl-model |
opus | Claude model for implementation |
--rate-limit-delay |
5.0 | Seconds between API calls |
