Research · SFT on legacy hardware
MalxLabs-V4
Proof that the “reasoning era” isn't restricted to H100 clusters. A 1.5B distilled reasoning model, fine-tuned in the cloud and tuned to think at 80+ tokens per second on a 2014-era GTX 1080.
- Parameters
- 1.5B
- Local speed
- 80+ tokens/s
- Recommended context
- 4,096 tokens
- Last commit
- —
- GitHub stars
- —
Fine-tune in the cloud. Run on hardware from 2014.
Through precise cloud fine-tuning on an RTX 4500 and aggressive inference tuning on a GTX 1080, legacy systems can perform high-level cognitive tasks reliably — where Project Kestrel trains a tiny architecture from scratch, this takes an existing model and makes it better at reasoning.
Cloud SFT
Supervised fine-tuning on math and code reasoning data, run on a rented 24 GB Ada GPU.
Local inference tuning
Full GPU offload, a deliberately short context, and a sampler tuned to stop a small model from rambling.
Measured, not claimed
Real head-to-head runs against the untuned base model on the same hardware, plus a non-GPU laptop benchmark.
Specification
What it is.
- Model name
- MalxLabs-V4
- Base model
- DeepSeek-R1-Distill-Qwen-1.5B
- Parameters
- 1.5 billion
- Quantization
- GGUF Q4_K_M
- Primary focus
- Technical reasoning and code generation
Training
How it was made.
- Hardware
- NVIDIA RTX 4500 Ada (24 GB) via RunPod
- Storage
- 80 GB volume, with data streaming for throughput
- Method
- Supervised fine-tuning (SFT)
| Dataset | Purpose |
|---|---|
| OpenMathInstruct-2 | Step-by-step mathematical reasoning |
| CodeFeedback-Filtered | Luau and Python script optimization |
Benchmarks: does the fine-tune actually help?
The repo includes a real head-to-head against the untuned base model — same hardware, same settings, back to back — so you can judge it rather than take a claim on faith. Third-party scores are cited from each model's own card; every MalxLabs-V4-vs-base number is a first-party measurement.
Read these as spot-checks, not leaderboards. The comparisons use small problem sets on one machine; the repo's full methodology, raw numbers and honesty notes on what is and isn't independently verified are in docs/HARDWARE_BENCHMARK_2026-08.md.
The legacy stack
Everything here is from 2014.
| Component | Specification |
|---|---|
| CPU | Intel i7-4790, overclocked to 4.0 GHz, core parking disabled |
| GPU | NVIDIA GTX 1080 — 8 GB VRAM, Pascal |
| RAM | 16 GB DDR3-1333 — the primary bottleneck |
| OS | Windows 10 ESU or Ubuntu 22.04 (WSL2) |
Performance: consistently 80+ tokens/s with all 35 layers offloaded to the GPU, which keeps long generations responsive.
Deployment config
The exact llama.cpp command.
cd ~/llama.cpp && ./build/bin/llama-cli \ -m ~/models/logic_model.gguf \ -ngl 35 \ -c 4096 \ --flash-attn on \ --mlock \ --temp 0.6 \ --min-p 0.1 \ --reasoning-format deepseek \ --jinja \ -p "<|User|>[INSERT PROMPT]<|Assistant|>"
What running it on old hardware taught.
The most useful results are the failure modes — and the small fixes that removed them.
~500 tokens of thought
A 1.5B model has a finite reach. It excels at short and medium chains but can fall into “hallucination loops” once its thinking runs past roughly 500 tokens.
DDR3 can't feed a big context
At 16k+ context the i7-4790 and DDR3 couldn't move the KV cache to the GPU fast enough, causing prefill hangs and terminal timeouts. Dropping to 4,096 tokens restored instant inference.
Min-P 0.1 fixes it
Early versions began answering inside the <think> block, caused by noisy low-probability tokens. Min-P 0.1 sampling muted them, forcing the model to finish reasoning before answering.
What it's good at
- Mathematical logic: step-by-step derivations for physics and algebra, most reliable on proofs and variable solving.
- Code generation: Luau (Roblox) and Python; best on specific functions rather than full applications.
- Reasoning depth: multi-step planning and chain-of-thought, logic puzzles and technical architecture questions.
Prompt with specifics
Extra specificity is the single biggest lever for code. For example:
Generate me a Python code that finds the average of 10 different numbers, the user inputs their desired numbers one by one, and the Python code will output the average of those 10 numbers.
The model returned a correct loop that reads ten inputs and prints their average — exactly the specification.
Try it in four steps.
- Have llama.cpp builtAnd an NVIDIA GPU with at least 8 GB of VRAM (a GTX 1080 or better).
- Download the GGUFThe Q4_K_M build from Hugging Face.
- Place it at
~/models/logic_model.ggufOr adjust the-mpath in the command above. - Run the deployment commandKeep context at or below 4,096 tokens and keep Min-P at 0.1.
Good applications
- Mathematical problem solving — physics and algebra derivations
- Targeted code generation in Luau and Python
- Multi-step logic puzzles
- Technical and architectural reasoning questions
Known limits: long chains of thought (past ~500 tokens) can loop, and large context windows stall on DDR3-class memory. Keep prompts specific and context short.
Recent commits.
Live from GitHub.
Commit history is loading — or see it on GitHub.