Research · from-scratch small LLM
Project Kestrel
A small, dense, adaptively-learning LLM architecture. Kestrel-Nano (~157M parameters) is trained and running — a from-scratch model built for coding, CLI and terminal work, tool use and secure coding, small enough to run on nearly any computer.
- Kestrel-Nano v0.1
- 157M params
- v0.1 training run
- 1.311B tokens
- v0.1 cost
- ~$21 (one rented RTX 3090)
- Last commit
- —
- GitHub stars
- —
The thesis
A small model that adapts in place, instead of a frozen giant.
A kestrel is a small falcon that hunts by hovering — reading the wind and adjusting continuously. That is the design thesis: stay light, but keep learning.
Not to be confused with Kestrel 1. Kestrel 1 is the self-hosted chat endpoint, which serves open-source models today. Project Kestrel is the from-scratch research program on this page.
Independence from the cloud
People in low-connectivity regions — and anyone who values privacy and ownership — deserve capable local models. Nano is ~150M parameters (~100 MB quantized): an 8 GB GPU, a CPU-only laptop, in principle a phone.
Small doesn't mean shallow
Recent evidence from looped models, hybrid attention and small-model research shows the gap closes through architecture shape, data quality and adaptive compute — not just scale.
The “frozen weights” critique
Today's models stop learning the moment training ends. Kestrel's memory subsystem — recurrent session state, sparse memory layers, nightly consolidation — is a direct engineering response.
Five pillars, each required to beat a plain baseline.
Every deviation from a vanilla transformer has to win against the matched-compute baseline — or it gets cut.
Kestrel-Nano · block stack (schematic)
Hybrid sequence mixer
Three gated linear-attention blocks to one full-attention block — O(1) inference state per token, long context on 8 GB, a persistent “mind-state.”
Evidence: Qwen3-Next, Kimi Linear, Samba
Looped core
The middle block group re-runs one to four times — deeper effective reasoning without more parameters, and compute that adapts to the problem.
Evidence: Ouro (1.4B ≈ 4B dense), Huginn
Product-key memory layers
Large sparse key-value banks give knowledge capacity without extra FLOPs — the “plastic tissue” that makes continual learning possible.
Evidence: Meta Memory Layers at Scale; Sparse Memory Finetuning (11% forgetting vs 89% for full fine-tuning)
Token efficiency
A superword BPE tokenizer plus a multi-token-prediction auxiliary head — roughly 20–30% more text per FLOP and per context window.
Evidence: SuperBPE (COLM 2025)
Layered memory & consolidation
Context flows into recurrent state, into memory slots, into nightly sparse consolidation — a model that learns from use without catastrophic forgetting.
Evidence: Titans, SEAL, Meta sparse-memory fine-tuning
What v0.1 actually showed — honestly.
Kestrel-Nano v0.1 generates coherent text and runs locally. The evaluation is deliberately unflattering: it's a record of what worked and what didn't.
157M params · 1.311B tokens
bf16 on a rented RTX 3090 for roughly $21, ending at a validation loss of 2.310.
Product-key memory
The memory gate validated — it converged to about 1.0, so the layers are genuinely being used.
The looped core
It changes the computation, but it isn't yet a quality dial you can turn for better answers.
Learned form, not content
The model is undertrained: it picked up how text looks before what it says — which is why the next phases focus on data and signal per token.
Full write-up: docs/09 — Nano v0.1 findings. A phase that finishes is compared against the v0.1 control before any new claim is made.
Live roadmap.
This checklist is read straight from the repository's README every few minutes — so it updates itself as the project moves, with no one editing this page.
The live roadmap couldn't be loaded right now. The same checklist is under “Status” in the repository README.
Read the reasoning.
The docs are the project's memory — including what was rejected and why.
| Doc | What's inside |
|---|---|
| MASTER-PLAN | The self-contained plan: names (Kestrel / Fledge / Roost), how the architecture and training work, phases, honest risks |
| 01 Goals & constraints | Mission, hard constraints, success criteria, and an honesty section |
| 02 Research survey | The annotated, linked study that backs every design choice |
| 04 Architecture spec | Blocks, presets (Nano / Mini) and parameter budgets |
| 05 Training methodology | Scale ladder, optimizer, data curriculum, continual-learning protocol |
| 06 Evaluation plan | Probes, benchmarks, baselines and “teach-me-today” tests |
| 08 Real-model roadmap | The forward plan: the local-only decision, Phase D tuning, what was rejected |
| 09 Nano v0.1 findings | The honest v0.1 evaluation |
| 10 Phase C plan | The 4.86B-token corpus and the dual-context v0.1 baseline |
| 11 v2 plan | Multi-token prediction, then distillation from a local teacher — buying quality with signal per token, not wall-clock |
What's in the code
kestrel/- The model and trainer: GLA hybrid, looped core, product-key memory, Muon + WSD optimizer, resumable, live status
kestrel/roost/- Lifelong learning: episodic store, session-state save/restore, and the eval-gated nightly consolidation job
scripts/- Corpus and SFT builders, throughput and MFU profiling, the generalization eval “scoreboard”, stopping-point fitting
tools/KestrelMonitor- A WPF live monitor with charts, probes, and Pause / Resume / Stop
Try the trained model
The 1.4 GB checkpoint and tokenized corpora exceed GitHub's limits and are kept locally — the corpora are reproducible with the repo's scripts, and the tokenizer is committed, so a checkpoint can be loaded and run interactively:
python -m kestrel.generate --interactive
Ground rules
Baked into every decision.
Portable by construction
Pure PyTorch, matmul-heavy, no custom kernels — written FP32/Pascal-safe for a GTX 1080, now running bf16 on an RTX 3060 with the same code.
Evidence or ablation
Every deviation from a vanilla transformer must beat the matched-compute baseline, or it gets cut.
Honest scale & budget
Everything runs on one RTX 3060 for the cost of electricity, and every training rung resumes the last — no work is ever thrown away.
Recent commits.
The project is worked on in the open — here's the latest, live from GitHub.
Commit history is loading — or see it on GitHub.