Field report

Teaching a 1.5B model to generate bash commands at gpt-4o level using 400k synthetic examples

How it started

Despite using LLMs for most of the coding, there was always one thing I kept Googling: Bash commands. It's quite flow-breaking to pause work, open Google, type the full query, go to Stack Overflow or similar, and look up the syntax I wanted.

Intuitively, this always felt like something a small model would perform well at because you have a finite set of commands, well-defined syntax, and easy-to-generate training data. So one day I decided to actually find out. Over the course of the experiment, I tried six different models: SmolLM 135M, SmolLM 360M, Qwen3-0.6B, Qwen2.5-Coder-1.5B, Qwen3.5-0.8B, and Qwen3.5-2B. The 0.6B and 1.5B models showed the best performance, so those checkpoints were preserved.

The good stuff you probably want

I mentioned the experiment on Hacker News, and people asked for the models, training data, and a write-up.

What started as a syntax reminder turned into an interesting exercise in teaching small models—and figuring out whether they were actually improving.

Astra ran the training loop

When I say I trained these models, I really mean Astra was the training controller, with different LLMs used to generate the synthetic dataset and another set of agents used to review and admit rows into it. Astra handled the training-and-evaluation workflow; I supplied the goal, constraints, and course corrections.

The starting point was coming up with a category of commands in a separate conversation and building a harness to generate and validate those commands. Initially we had about 30k samples. Then the loop was simple: train, evaluate (I got Astra to create some benchmarks), analyze, generate more data, brainstorm different ways to train, and run the next training runs. The whole process took about 10 days, most of it spent with models working on their specific tasks: generate, review, evaluate.

The prompt was some variation of (among other things):

We only care about English-to-Bash commands, nothing else. As long as we make gains in that domain, we are happy to sacrifice any abilities in any other domain. We want the best goddamn English-to-Bash model ever created in the history of humanity, do what you need to do.

The only real direction I provided was keeping a mental map of multiple trained checkpoints, prompting the models to compare those in specific categories, asking them to investigate why a checkpoint traded gains in one area for losses in another, and so on.

The released models use supervised fine-tuning with LoRA, starting from Qwen3-0.6B and Qwen2.5-Coder-1.5B-Instruct. The 4-bit versions surprisingly keep most of the gains of the 16-bit versions, and they are convenient to run on CPU systems.

More examples weren't always the answer

The published dataset contains 401,975 distinct request/answer pairs, covering file operations, text processing, Git, archives, networking, quoting, and composed commands. A row looks like this:

{"request":"list files in this directory","response":{"kind":"COMMAND","value":"ls"}}

Examples were synthetically generated, reviewed, and checked through execution where available. Not every row was individually executed. Different descriptions of the same command were intentionally retained.

One experiment made the difference between learning examples and learning behavior very clear. A 0.6B model improved from 8/72 to 64/72 on time-related training requests. On separate time-transfer requests, it stayed at 1/18. Meanwhile, routine-command performance fell from 70/102 to 60/102.

The model was learning the supplied answers. That wasn't the same as reliably applying them elsewhere.

Targeted repairs brought another problem: retention. A later 1.5B checkpoint passed five more ALFA-updated cases than the model I released, but lost 63 passes on the internal suite. I kept the more balanced checkpoint. Exporting also changed checkpoint rankings, so evaluating the actual quantized file became part of selection rather than an afterthought.

The detailed results preserve those comparisons, including the experiments that didn't help.

I ended up debugging the benchmark too

I used ALFA, a 300-task shell benchmark, alongside internal execution tests. Investigating disagreements uncovered both model mistakes and evaluator problems: argument transport, fixture resets, reference commands, and checks that could accept wrong outputs or reject correct alternatives.

That work became ALFA-updated. It keeps the task set but documents the repaired environment and grading rules. It's a separate benchmark definition, not a silent replacement for upstream ALFA.

Benchmark comparison

The table combines our local results with scores reported by the whatisit project and the barbarabhb model card. ALFA-updated? identifies the benchmark used: Yes for our repaired version, No for original ALFA. whatisit appears twice: its published original-ALFA score was 62.0%, while our ALFA-updated run measured 63.7%—a 1.7-percentage-point increase in the reported score.

Shell-command benchmark results, with benchmark version and source
Model / configuration Size Pass rate ALFA-updated? Reported by
GPT-4o (published) Cloud API 73.0% No ALFA authors
EasyCommand 1.5B, Q4_K_M 940.4 MiB 212/300 — 70.7% Yes Our local run
nl2sh-3b 1.9 GB 65.7% No whatisit
barbarabhb / nl2sh-qwen25-coder-1.5b, Q4_K_M 941 MB 65.67% No barbarabhb
barbarabhb / nl2sh-qwen25-coder-1.5b, Q4_K_M + imatrix 941 MB 65.00% No barbarabhb
barbarabhb / nl2sh-qwen25-coder-1.5b, Q6_K 1.2 GB 64.33% No barbarabhb
barbarabhb / nl2sh-qwen25-coder-1.5b, Q8_0 1.6 GB 63.67% No barbarabhb
whatisit / nl2sh-1.5b, Q4_K_M 941 MB (reported) 191/300 — 63.7% Yes Our local run
whatisit / nl2sh-1.5b, Q4_K_M 941 MB 62.0% No whatisit
Qwen2.5-Coder-7B, untuned 4.4 GB 61.3% No whatisit
EasyCommand 0.6B, Q8_0 767.5 MiB 174/300 — 58.0% Yes Our local run
EasyCommand 0.6B, Q4_K_M 461.8 MiB 165/300 — 55.0% Yes Our local run
Qwen2.5-Coder-1.5B, untuned 941 MB 54.0% No whatisit

21 more passes than whatisit (+7 points) on ALFA-updated, using each tool's own settings (256-token limit for EC, 64 for whatisit).

A specialist peaked at 217/300 (72.3%), but lost 63 internal-test passes. I shipped the more balanced checkpoint.

These tests informed development, not an independent holdout. GPT-4o's published 73–74% used original ALFA; similar scores don't establish parity. Full methodology and run details are in the evaluation notes.

Try it, or keep training

If you want to run it locally, ec embeds llama.cpp and offers command preview and interactive execution. Start with --preview; a plausible command can still be wrong. The target is GNU/Linux Bash, not every shell or platform.

If you want to experiment, the reusable pieces are:

Models and data are Apache-2.0; application and benchmark code are MIT. The flat dataset doesn't reproduce historical weighting, and it has no official test split. If you train with it, separate holdouts by task family rather than randomly splitting paraphrases.

I'd like to see what others can improve: better coverage, more reliable transfer, or fresh tests that expose something I missed. This started with wanting a command reminder. Sharing the materials seems like a good way to keep learning.