Introducing synthlite: prompts in, a fine-tuning dataset out
The open-source Rust CLI we use to build training data for Sia, our IT-operations agent.

At Scogo AI we're building Sia, an autonomous agent for IT operations. Fine-tuning a model for that work takes a lot of domain-specific supervised data. This post covers the tool we developed internally to make that data, the choices behind it, and what each choice cost us. The tool is synthlite, a Rust CLI, and it's now open source under Apache-2.0.
The problem
We had prompts. We needed answers to them in a format our training stack reads, and we needed that over and over: new domains, new teacher models, new prompt sets.
We literally looked at 23 open-source synthetic-data tools; the README compares synthlite with the ten best known. Most are Python frameworks or apps that also write the prompts, run multi-step pipelines, or offer model-based scoring and judge steps. Those steps can help, but they add calls per row, add latency, increased cost and bring in a second model whose mistakes you also have to audit. None of the 23 was a compiled, judge-free generator that makes one call per prompt you already have.
We wanted something minimalistic. Take a file of prompts, get one answer per prompt from a teacher, drop the obviously broken answers with rules we can read, and hand back a dataset we then evaluate ourselves. It also had to survive being killed at any point without duplicating rows or paying for them twice.
Design principles, and what each one cost
One call per prompt
Each prompt gets exactly one teacher call. Retries happen only on transient errors such as timeouts and 5xx responses. There's no best-of-K and no judge. That makes spend predictable: about one request per prompt, with hard caps from --max-requests and --max-rows, and --dry-run shows the plan without calling anything.
What we gave up: selection. A judge or best-of-K can catch a fluent wrong answer. We can't. That makes the teacher the quality ceiling, so for datasets we train on we use a frontier model, and we recommend the same: Claude Opus 5.5 or GPT-6-Sol.
Seeds in, dataset out
synthlite doesn't write, paraphrase or evolve prompts. You bring a .txt file with one prompt per line, JSONL records with prompt, id and metadata, or records from Taskgen, our seed generator.
What we gave up: help with diversity. If your seed set is narrow or repetitive, the dataset will be too.
A deterministic gate
synthlite gate runs checks with no model involved: exact dedup of normalized prompts and then responses, length bounds, a refusal-phrase check on the opening of each reply, repeated-line detection to catch loops, and a list of held-out prompt hashes so eval prompts can't leak into training. Every rejected row is written to rejected.jsonl with its reason, and the same input always gives the same output.
What we gave up: semantic checks. The gate can tell that a reply is too short or stuck in a loop. It can't tell whether the reply is right.
The committed row is the checkpoint
There's no separate checkpoint file. A single writer appends each row to rows.jsonl and fsyncs it before counting it as committed. On startup, a torn last line left by a crash mid-write is copied to a quarantine file and cut off. Kill the process at any point and run the same command again. It picks up where it stopped, and committed rows are never requested again.
What we gave up: batched writes. One fsync per row is slower than group commit. But a provider call takes seconds and an fsync takes milliseconds, so the disk isn't the bottleneck.
Content-addressed identity
Row ids are hashes of the prompt record. The train/validation split is by id, so a prompt lands in the same split on every run and every machine. A hash of the generation config is stored in the output folder, and a run with a different config refuses to write there, so two configs never mix in one dataset.
What we gave up: changing settings mid-run. Change the temperature or the system message and you need a new output folder.
One static binary
synthlite is a Rust binary of about 6–7 MB: static musl builds for Linux x86_64 and aarch64, plus macOS arm64 and x86_64. There's no Python environment to set up and no service to run. You can install it with a one-line script, with cargo install synthlite, or from the https://github.com/scogo-ai/synthlite/pkgs/container/synthlite image.
What we gave up: Python extensibility. You can't drop a custom filter function into the gate. You post-process the JSONL instead.
Any OpenAI-compatible endpoint
It works with any OpenAI compatible endpoint that serves the chat-completions API. Several keys can be pooled. Concurrency per key starts at 16, halves on a 429 and grows back one slot at a time after clean answers, and Retry-After is respected. A hosted API and a single local GPU both work without tuning.
Private by default
synthlite push uploads to a private Hugging Face dataset and refuses to push to a public one. Our demo datasets are public only because we changed their visibility by hand after pushing.
Training-shaped output
Rows use the standard chat messages format (system, user, assistant), which TRL, Axolotl, Unsloth and Hugging Face datasets load as-is. Each row also carries lineage metadata: the teacher's provider, base URL and model, the model that actually answered, token usage and timestamps.
What we gave up: other formats. It's single-turn SFT only. There are no preference pairs for DPO and no multi-turn conversations yet.
Decision traces
For an IT-operations agent we also wanted answers that show how they got there. --detailed keeps one call per prompt but asks the teacher for a JSON object of typed steps (Evidence, Hypothesis, Action, Verification, Conclusion) followed by a final answer. synthlite validates the structure and renders it as numbered steps. A reply that isn't a valid trace is recorded as invalid_trace and never becomes a training row. With small models this matters: in our demo run, 16 of 25 traces were valid on the first pass, and --retry-failed recovered the rest in two rounds.
[generation].persona names the assistant in the system message. Ours is "Sia by Scogo.AI, which delivers Autonomous Agentic IT Operations". The trace instructions were written for IT operations, and the validation checks that the steps are present, not that they're correct. There's a sample at https://huggingface.co/datasets/ScogoAI/synthlite-demo-itops-decision-traces.
Try it in four commands
You need an OpenRouter key, a Hugging Face write token and a prompts.txt with one prompt per line. To try it without spending anything, this uses OpenRouter's free router, openrouter/free, which picks a free model for each request; synthlite records the model that answered in metadata.served_model.
curl -fsSL https://raw.githubusercontent.com/scogo-ai/synthlite/main/install.sh | shcurl -fsSLO https://raw.githubusercontent.com/scogo-ai/synthlite/main/examples/configs/openrouter-free.tomlOPENROUTER_API_KEY=sk-or-... synthlite prompts.txt --config openrouter-free.tomlHF_TOKEN=hf_... synthlite push --hf-repo your-name/my-first-sft
push runs the gate first if you haven't run it yet. Our demo run made 25 rows from 25 prompts in about 5 minutes, answered by 9 different models: https://huggingface.co/datasets/ScogoAI/synthlite-demo-itops-sft. For a dataset you'll train on, change model to your frontier teacher.
Limitations
Without a judge, the output is unverified teacher text. Evaluate before you train.
Quality depends on the teacher, hence always use SOTA model. With a router, filter on metadata.served_model if one model underperforms.
It supports single-turn SFT only: no DPO pairs and no multi-turn conversations.
It works with OpenAI-compatible APIs only.
--detailed is tuned for IT operations. We haven't tuned its instructions for other domains.
Roadmap
Taskgen: our seed generator, which writes the structured prompts we feed to synthlite. We plan to open-source it as well.
Presets: ready-made configurations for common setups.
Export formats: more options beyond chat messages JSONL.
A public benchmark: synthlite's single call and deterministic gate against a judge pipeline, with the method and data published. We don't have results yet and won't claim any until we do.
Try it and tell us
The code is at https://github.com/scogo-ai/synthlite, with 120+ automated tests that run against fake provider and Hugging Face servers. The feedback we want most: rows that got through the gate and shouldn't have, and the export format you need next. Issues and pull requests are welcome.
Written by
Published on


