Research · from-scratch small LLM

Project Kestrel

A small, dense, adaptively-learning LLM architecture. Kestrel-Nano (~157M parameters) is trained and running — a from-scratch model built for coding, CLI and terminal work, tool use and secure coding, small enough to run on nearly any computer.

Pure PyTorch Trained on one GPU Hybrid attention · looped core · sparse memory
Abstract flowing artwork for Project Kestrel.
Kestrel-Nano v0.1
157M params
v0.1 training run
1.311B tokens
v0.1 cost
~$21 (one rented RTX 3090)
Last commit
—
GitHub stars
—

The thesis

A small model that adapts in place, instead of a frozen giant.

A kestrel is a small falcon that hunts by hovering — reading the wind and adjusting continuously. That is the design thesis: stay light, but keep learning.

Not to be confused with Kestrel 1. Kestrel 1 is the self-hosted chat endpoint, which serves open-source models today. Project Kestrel is the from-scratch research program on this page.

01

Independence from the cloud

People in low-connectivity regions — and anyone who values privacy and ownership — deserve capable local models. Nano is ~150M parameters (~100 MB quantized): an 8 GB GPU, a CPU-only laptop, in principle a phone.

02

Small doesn't mean shallow

Recent evidence from looped models, hybrid attention and small-model research shows the gap closes through architecture shape, data quality and adaptive compute — not just scale.

03

The “frozen weights” critique

Today's models stop learning the moment training ends. Kestrel's memory subsystem — recurrent session state, sparse memory layers, nightly consolidation — is a direct engineering response.

Five pillars, each required to beat a plain baseline.

Every deviation from a vanilla transformer has to win against the matched-compute baseline — or it gets cut.

Kestrel-Nano · block stack (schematic)

15 gated linear-attention blocks — O(1) state per token, no positional encoding 5 full-attention blocks — the only ones carrying RoPE
Context window Recurrent session state Sparse memory slots Nightly consolidation
Simplified schematic. It shows the 3 : 1 mix and the four-tier memory ladder; ordering within the stack is illustrative. The exact block layout, looped-core placement and memory-layer positions are in docs/04-architecture-spec.md.
01

Hybrid sequence mixer

Three gated linear-attention blocks to one full-attention block — O(1) inference state per token, long context on 8 GB, a persistent “mind-state.”

Evidence: Qwen3-Next, Kimi Linear, Samba

02

Looped core

The middle block group re-runs one to four times — deeper effective reasoning without more parameters, and compute that adapts to the problem.

Evidence: Ouro (1.4B ≈ 4B dense), Huginn

03

Product-key memory layers

Large sparse key-value banks give knowledge capacity without extra FLOPs — the “plastic tissue” that makes continual learning possible.

Evidence: Meta Memory Layers at Scale; Sparse Memory Finetuning (11% forgetting vs 89% for full fine-tuning)

04

Token efficiency

A superword BPE tokenizer plus a multi-token-prediction auxiliary head — roughly 20–30% more text per FLOP and per context window.

Evidence: SuperBPE (COLM 2025)

05

Layered memory & consolidation

Context flows into recurrent state, into memory slots, into nightly sparse consolidation — a model that learns from use without catastrophic forgetting.

Evidence: Titans, SEAL, Meta sparse-memory fine-tuning

What v0.1 actually showed — honestly.

Kestrel-Nano v0.1 generates coherent text and runs locally. The evaluation is deliberately unflattering: it's a record of what worked and what didn't.

TRAINING RUN

157M params · 1.311B tokens

bf16 on a rented RTX 3090 for roughly $21, ending at a validation loss of 2.310.

WORKED

Product-key memory

The memory gate validated — it converged to about 1.0, so the layers are genuinely being used.

NOT YET

The looped core

It changes the computation, but it isn't yet a quality dial you can turn for better answers.

THE VERDICT

Learned form, not content

The model is undertrained: it picked up how text looks before what it says — which is why the next phases focus on data and signal per token.

Full write-up: docs/09 — Nano v0.1 findings. A phase that finishes is compared against the v0.1 control before any new claim is made.

Live roadmap.

This checklist is read straight from the repository's README every few minutes — so it updates itself as the project moves, with no one editing this page.

The live roadmap couldn't be loaded right now. The same checklist is under “Status” in the repository README.

Read the reasoning.

The docs are the project's memory — including what was rejected and why.

DocWhat's inside
MASTER-PLANThe self-contained plan: names (Kestrel / Fledge / Roost), how the architecture and training work, phases, honest risks
01 Goals & constraintsMission, hard constraints, success criteria, and an honesty section
02 Research surveyThe annotated, linked study that backs every design choice
04 Architecture specBlocks, presets (Nano / Mini) and parameter budgets
05 Training methodologyScale ladder, optimizer, data curriculum, continual-learning protocol
06 Evaluation planProbes, benchmarks, baselines and “teach-me-today” tests
08 Real-model roadmapThe forward plan: the local-only decision, Phase D tuning, what was rejected
09 Nano v0.1 findingsThe honest v0.1 evaluation
10 Phase C planThe 4.86B-token corpus and the dual-context v0.1 baseline
11 v2 planMulti-token prediction, then distillation from a local teacher — buying quality with signal per token, not wall-clock

What's in the code

kestrel/
The model and trainer: GLA hybrid, looped core, product-key memory, Muon + WSD optimizer, resumable, live status
kestrel/roost/
Lifelong learning: episodic store, session-state save/restore, and the eval-gated nightly consolidation job
scripts/
Corpus and SFT builders, throughput and MFU profiling, the generalization eval “scoreboard”, stopping-point fitting
tools/KestrelMonitor
A WPF live monitor with charts, probes, and Pause / Resume / Stop

Try the trained model

The 1.4 GB checkpoint and tokenized corpora exceed GitHub's limits and are kept locally — the corpora are reproducible with the repo's scripts, and the tokenizer is committed, so a checkpoint can be loaded and run interactively:

Shell
python -m kestrel.generate --interactive

Ground rules

Baked into every decision.

01

Portable by construction

Pure PyTorch, matmul-heavy, no custom kernels — written FP32/Pascal-safe for a GTX 1080, now running bf16 on an RTX 3060 with the same code.

02

Evidence or ablation

Every deviation from a vanilla transformer must beat the matched-compute baseline, or it gets cut.

03

Honest scale & budget

Everything runs on one RTX 3060 for the cost of electricity, and every training rung resumes the last — no work is ever thrown away.

Recent commits.

The project is worked on in the open — here's the latest, live from GitHub.

Commit history is loading — or see it on GitHub.