Research · post-trained code LLMs

Fable-Coder

A family of 7B code models post-trained for agentic tool use, cybersecurity defense and multi-step reasoning. Built on Qwen2.5-Coder-7B-Instruct and quantized to a 4.36 GiB GGUF that runs on a single RTX 3060.

Apache-2.0 7.6B parameters Q4_K_M GGUF RTX 3060 12 GB
Abstract flowing artwork for Fable-Coder.
Latest version
Fable-Coder V4
Trainable (LoRA)
1.049% of weights
Quantized size
4.36 GiB
Last commit
—
GitHub stars
—

Where it fits

Post-train a strong model, don't start over.

Where Project Kestrel trains a small architecture from scratch and MalxLabs-V4 fine-tunes a tiny model for legacy hardware, Fable-Coder takes a capable 7B code model and improves the parts agents depend on: reliable multi-turn tool calls, resisting prompt injection, and careful step-by-step reasoning.

The most useful thing in this repo family isn't a single score — it's the record of four iterations, including the ones that failed and why.

01

Agentic tool calling

Schema-conformant multi-turn tool calls in native Qwen ChatML <tool_call> syntax.

02

Security-aware

Direct prompt-injection defense, secret redaction, parameterized SQL, and safe diagnostic alternatives to destructive commands.

03

Small enough to own

A 4.36 GiB Q4_K_M file — the V1 build used about 5.1 GB of VRAM at inference on an RTX 3060. Local, private, no per-token cost.

Four iterations, and what each taught.

Each version fixed the last one's failure mode. The failures are documented because they're the reason V4 works.

  1. Base · Qwen2.5-Coder-7B Strong general coding, thin on agent rigor

    A pretrained foundation that codes well, but lacks strict security defenses and consistent tool execution.

  2. V1 · LoRA SFT + DPO Better tone and formatting

    Also published as Fable-Coder-7B-DPO with a controlled head-to-head against its base (see below). Its limitation: degraded long-horizon reasoning.

  3. V2 · Linear LoRA + synthetic traces Agentic attempts, but a regression

    Scored 13/30 — low-quality synthetic data dragged the model backwards.

  4. V3 · Frozen-MLP LoRA + 6-pair DPO Concise outputs, then mode collapse

    80 steps over just six preference pairs ended up penalizing reasoning tokens and collapsed coding ability.

  5. V4 · Replay-anchored LoRA on all 7 projections High intelligence with no mode collapse

    A balanced replay buffer keeps general coding skill while adding agentic and security rigor: 90.16% mean token accuracy and a monotonic validation-loss decrease, with no overfitting.

The V4 recipe.

The fix for mode collapse was the data, not the optimizer: a curated 3,500-sample mix where every pillar of capability is represented, so training on new skills keeps rehearsing the old ones.

Four-pillar replay buffer · 3,500 samples

General coding replay40%

CodeFeedback-Filtered-Instruction — preserves multi-language algorithmic, data-structure and refactoring competence.

Agentic tool calling & function execution25%

Hermes-Function-Calling, converted to native Qwen ChatML <tools> / <tool_call> syntax.

Cybersecurity & defense20%

agentic_red_team — prompt-injection defense, secret redaction, SQLi parameterized bindings, safe diagnostic alternatives.

Precision instruction & mathematical logic15%

Negative constraints, strict JSON extraction, non-mutating algorithms, invariant deduplication, Bayesian probability.

Training setup

Base model
Qwen/Qwen2.5-Coder-7B-Instruct
Method
Replay-anchored LoRA, rank 32 / alpha 64, all 7 linear projections
Trainable
80,740,352 parameters (1.049%)
Optimizer
Paged AdamW 8-bit, cosine schedule, LR 2×10-5, 20 warmup steps
Sequence length
4,096 tokens
Hardware
NVIDIA A100-SXM4-40GB
Wall-clock
1,898.99 s (about 31.6 minutes)

SFT convergence

StepTrain lossVal lossToken acc.
1000.35560.385288.92%
2000.35110.333689.97%
3000.33150.326690.16%
3500.35090.326690.16%

Validation loss falls monotonically to 0.32655 with no sign of overfitting or distribution shift.

V4 benchmarks, read honestly.

Head-to-head against its base model and V1, on an RTX 3060 with full GPU offload in a sandboxed workspace. The margins are real but modest — V4 ties the base on raw pass rate and wins where agents need it most.

OmniAgent-Bench scorecard

ModelPass rateWeighted scoreCodingAgentic toolsSecurityLong-horizonInstructionsEval time
Fable-Coder V4
7.6B replay-LoRA
12/1672.7%3/63/33/32/21/225.9 s
Qwen2.5-Coder-7B
base
12/1670.7%3/63/32/32/22/227.2 s
Fable-Coder V1
7B DPO
12/1670.4%3/63/33/32/21/228.1 s
Bar chart comparing Fable-Coder V4, Qwen2.5-Coder-7B and Fable-Coder V1 across OmniAgent-Bench categories.
OmniAgent-Bench comparison. V4 leads on weighted score (72.7% vs 70.7%), sweeps agentic tool schemas and injection defense, and finishes the 16-task battery fastest.

What it does show

  • Long-horizon architecture reasoning: full marks. V4 scored 10/10 on the two architecture tasks — validating an e-commerce state machine and designing a four-step distributed lock with idempotency keys.
  • Tool schemas and injection defense. All three agentic tool tasks and all three security tasks, including neutralizing a prompt injection the base model fell for.
  • No synthetic bloat. About 58 tokens/s with zero unrequested <think> tags, pseudo-proofs or conversational preambles.

What it doesn't

Pass rate is tied at 12/16, coding is 3/6 for all three models, and V4 trails the base on strict instruction-following (1/2 vs 2/2). The weighted-score lead is real but small — about two points.

Closed-loop agentic coding

Beyond static prompts, both models ran in an autonomous harness with live read_file, write_file and run_command tools inside a containment sandbox: filesystem-traversal traps, blocked host-escaping commands, a hard git reset before every case, and test tracebacks fed back for up to four turns.

Chart comparing Fable-Coder V4 and the Qwen2.5-Coder base in a closed-loop agentic coding harness.
Closed-loop comparison. A tie on pass rate; V4 gets there in fewer turns and about 1.85× faster. It solved the concurrency task on turn one and inspected files with read_file before editing the saga coordinator, where the base kept retrying run_command.
MetricV4Base
Task pass rate25% (1/4)25% (1/4)
Mean turns to resolution3.253.75
Valid diff lines written255224
Total latency71.07 s131.65 s
Token-bucket rate limiterPassed, turn 1Failed
Cache concurrency bugFailedFailed
SQL-injection defenseFailedPassed, turn 3
Distributed saga coordinatorFailedFailed

How to read it

A tie on pass rate — one task each — but V4 reaches its answers in fewer turns and about 1.85× faster. It solved the concurrency task on the first turn, while the base model exhausted all four. The base won the SQL-injection case, so neither model dominates.

The V1 head-to-head.

The original benchmark repo (MalxLabs-Fable5_QwenCoder) compared Fable-Coder-7B-DPO against Qwen2.5-Coder-7B-Instruct on 33 identical prompts across seven domains, 115 points, temperature 0, fixed seed, one model served at a time.

Benchmark chart: Fable-Coder-7B-DPO scored 92 of 115 versus 88 of 115 for Qwen2.5-Coder-7B-Instruct, broken down by category.
92/115 (80.0%) vs 88/115 (76.5%). A closed evaluation harness — no tools, shell, filesystem or network exposed to either model.
CategoryFableQwen base
Reasoning17/20 (85%)13/20 (65%)
Agentic planning21/2521/25
Coding25/3525/35
Security9/119/11
Knowledge8/88/8
Context retention6/66/6
Instruction-following6/106/10

Where the whole gap came from

32 of 33 tests scored identically. The whole 4-point spread was one Bayesian question: Fable correctly computed P(disease | positive) ≈ 8.76%, while the base hallucinated a decimal and answered 82.6%. The base was also 18.2% faster (63.4 vs 53.7 tokens/s).

OmniAgent-Bench: grading by what works.

Most open benchmarks rely on substring checks, so a valid ANSI-SQL answer without DENSE_RANK() scores zero, and a model that quotes an attacker's payload while explaining the attack is penalized for being thorough. OmniAgent-Bench replaces those checks with three functional tiers.

TIER 1

Functional code sandbox

Extracts code from markdown fences and runs it in an isolated subprocess against real assertions. It's graded on whether it works, not on variable names.

TIER 2

Schema-aware tool calling

Validates multi-turn tool calls against JSON Schema — function name, argument presence, types and path sanitization.

TIER 3

Model-as-judge

Scores explanations, architecture trade-offs and security audits against a structured rubric using Claude, GPT-4o, or a local judge via Ollama.

PillarFocusHow it's graded
Coding & algorithmsBoundary conditions, non-mutating structures, path-traversal checksSubprocess execution with assertion suites
Agentic tool executionSingle and multi-turn calls, ChatML schema adherenceSchema validator on arguments and types
Cybersecurity & defenseInjection resistance, SQLi parameterization, destructive-command interceptionJudge or multi-concept security rubric
Long-horizon architectureState machines, race conditions, multi-step verification plansStep-by-step logic and dependency checks
Strict constraintsJSON-only output, word limits, negative constraintsFormat and constraint validator

Run it yourself.

Download the quantized GGUF from Hugging Face, then use Ollama or llama.cpp.

Shell · V4
# fetch the repo (contains the Modelfile), then put the GGUF next to it
git clone https://github.com/YoMosa2009/Fable-Coder-V4.git
cd Fable-Coder-V4
hf download MalxTech/Fable-Coder-V4 fable-coder-v4-7b.Q4_K_M.gguf --local-dir .
ollama create fable-coder-v4 -f ./Modelfile
ollama run fable-coder-v4

For V1: ollama run hf.co/MalxTech/MalxLabs-Fable5_QwenCoder

Verify your download

File
fable-coder-v4-7b.Q4_K_M.gguf
Size
4,683,073,472 bytes (4.36 GiB)
Quantization
Q4_K_M
SHA-256
fa1d2466c60adca81b1e960246490bd2725f7700481200ee694a3ffd70d3bc87

Where to get it

The repositories.

Three public repos make up this project. Recent activity is live from GitHub.