AgentsAutonomy & tool use
Hugging Face's ML Intern agent builds six niche models for about $103, including a 0.8B Qwen-Image prompt rewriter that runs on a CPU
Hugging Face staff used the ML Intern agent in HuggingChat to fine-tune and distill six small public models over a few days. The compute cost about $103 in total, and no single build cost more than $37.

The Pocket Rewriter is a 0.8B student model distilled from Qwen-Image 2.1's 9B prompt rewriter. It was trained on 8,797 requests labelled by that larger model and outputs compact JSON with no thinking tokens. On 300 held-out requests, 99.7% of its outputs were valid. It used about 453 tokens per request against the teacher's 1,631 and matched the teacher's roughly 53% fidelity on quoted text. On Arabic-script text, fidelity fell to 33.3%. [1] [2]
Each build was guided by a structured prompt that grew from about 450 words to about 2,000. Every prompt listed facts the agent should not re-derive, required a zero-shot baseline, and ran a 50-step smoke test to confirm the weights actually changed. A hard spending cap applied, and the agent asked the user before starting any paid job. [1]
Read the headline numbers alongside how they were measured. Agate 4-step scores 0.536 on GenEval against its 50-step teacher's 0.563, with overlapping confidence intervals, and runs about 12.8× faster. Citrus Doctor's strict accuracy rose from 14.9% to 52.8%, but its relaxed metric relies on an alias table that was tuned on the test set. [8] [9]
All figures come from the authors' own model cards; none has been tested independently. If they hold, a narrow, deployable model becomes mostly a prompt-writing and evaluation problem rather than a compute one.
Read the full assessment
That lowers the bar for teams that need specialized small models.
Executive brief
A Hugging Face developer used the company's ML Intern agent, which runs inside HuggingChat, to build six public models in a few days for about USD 103 in total compute. The most notable is a 0.8B "Pocket Rewriter" distilled from Qwen-Image 2.1's 9B prompt rewriter. It runs on a CPU and returns valid output 99.7% of the time, using about 28% of the teacher's tokens (HF blog; model card). Every result comes from the authors' own model cards. No outside group has checked them. Several of the models can't be used commercially because of their licenses.
What changed and event timeline
Ml-intern launches
Hugging Face released an open-source agent built on smolagents that automates LLM post-training, from reading the literature through to evaluation (;).
Qwen-Image 2.1 ships
Alibaba released a 7B DiT image model plus two 9B prompt-rewriter checkpoints (PE-T2I and PE-I2I). It also moved from Apache 2.0 to a non-commercial research license (;).
Standalone repo archived
The ml-intern GitHub repository (Apache-2.0, about 10.8k stars) was archived, and its functionality now lives in HuggingChat ().
Six-model case study
Yuvraj Sharma and Abubakar Abid published the six builds, with costs per build and all seven prompts on GitHub ().
Capabilities and access
- Access: HuggingChat with "ML-intern mode" switched on. Jobs run on Hugging Face hardware.
- Models built:
- Citrus Doctor (fine-tune of Qwen3.5-2B)
Read the full section
- Access: HuggingChat with "ML-intern mode" switched on. Jobs run on Hugging Face hardware. Each task starts with a zero-dollar budget, and the agent asks the user before running any paid job (HF blog).
- Models built:
- Citrus Doctor (fine-tune of Qwen3.5-2B)
- Huggy LoRA (FLUX.2 klein base 4B)
- Pocket Rewriter (0.8B and 2B students)
- Viewpoint Orbit and Doodle-in LoRAs (Qwen-Image 2.1)
- Agate 4-step, a distillation of the 260M-parameter Agate Preview 002 that also runs in the browser via ONNX
- Cost per build: USD 1.90 to USD 37 (HF blog).
Technical analysis
- Prompting method: prompts started around 450 words and grew to about 2,000.
- Pocket Rewriter recipe: 8,797 synthetic requests were labelled by the 9B teacher.
- Agate recipe: guidance was baked into the weights and the step count cut in stages (16 → 8 → 4).
Read the full section
- Prompting method: prompts started around 450 words and grew to about 2,000. Each one included a "Verified facts, do not re-derive" section, a required zero-shot baseline, a 50-step smoke test that checks the weights actually changed, and a hard spending cap (HF blog).
- Pocket Rewriter recipe: 8,797 synthetic requests were labelled by the 9B teacher. The student was trained with SFTTrainer for two epochs to output compact JSON with no thinking tokens, then tested on 300 held-out requests (model card).
- Agate recipe: guidance was baked into the weights and the step count cut in stages (16 → 8 → 4). The card reports paired comparisons on the same prompts and seeds, with confidence intervals (model card).
- Reproducibility: prompts, datasets and model cards are all public. The agent's runs, however, are not deterministic.
Claims and evidence
- Citrus Doctor: strict accuracy rose from 14.9% to 52.8% on 335 test images.
- Pocket Rewriter: 99.7% of outputs are valid JSON, averaging 453 tokens against the teacher's 1,631.
- Agate 4-step: GenEval score of 0.536 [0.508, 0.564], against the 50-step teacher's 0.563 [0.534, 0.591], and about 12.8× faster (model card).
Read the full section
- Citrus Doctor: strict accuracy rose from 14.9% to 52.8% on 335 test images. Relaxed accuracy rose from 20.3% to 61.2%, but the alias table behind the relaxed metric was tuned on the test set (model card).
- Pocket Rewriter: 99.7% of outputs are valid JSON, averaging 453 tokens against the teacher's 1,631. It matches the teacher's roughly 53% fidelity on quoted text. The blog gives 1,840 training examples; the card says 1,776 filtered pairs (model card).
- Agate 4-step: GenEval score of 0.536 [0.508, 0.564], against the 50-step teacher's 0.563 [0.534, 0.591], and about 12.8× faster (model card).
- Doodle-in: objects were detected where they were drawn 67.5% of the time. Results on unseen classes were similar (65.0% vs 64.2%) (HF blog).
- Who measured: the authors are Hugging Face staff. None of the six models has been evaluated independently.
Context and prior work
- The April launch result: ml-intern took Qwen3-1.7B from about 10% to 32% on GPQA, beating Claude Code's 22.99% (MarkTechPost). That task comes from PostTrainBench, built by researchers at Tübingen and the Max Planck Institute.
- What's new here: this post moves from that benchmark to small, specific models that nobody else had built. It's also the first showcase since the agent moved into HuggingChat (GitHub).
Read the full section
- The April launch result: ml-intern took Qwen3-1.7B from about 10% to 32% on GPQA, beating Claude Code's 22.99% (MarkTechPost). That task comes from PostTrainBench, built by researchers at Tübingen and the Max Planck Institute. The coverage doesn't show anyone reproducing the run independently (AgentConn).
- What's new here: this post moves from that benchmark to small, specific models that nobody else had built. It's also the first showcase since the agent moved into HuggingChat (GitHub).
Limitations, safety and contested findings
- Citrus Doctor: its treatment text is templated from a knowledge base, not written per image.
- Pocket Rewriter: fidelity on Arabic-script text is only 33.3%, against 57.1% for Latin script (model card).
- Agate 4-step: outputs are fixed at 256×256, it ignores negative prompts, and it handles exact text and negation poorly (model card).
Read the full section
- Citrus Doctor: its treatment text is templated from a knowledge base, not written per image. The card recommends review by an agronomist, notes that nutrient deficiencies are still poorly detected, and warns that results may drop on new varieties, backgrounds or cameras (model card).
- Pocket Rewriter: fidelity on Arabic-script text is only 33.3%, against 57.1% for Latin script (model card).
- Agate 4-step: outputs are fixed at 256×256, it ignores negative prompts, and it handles exact text and negation poorly (model card).
- Run reliability: the Viewpoint project needed 48 jobs, some of which failed on missing packages or wrong paths and had to be resubmitted (HF blog).
- Agent constraints: earlier reviews flagged lock-in to the Hugging Face ecosystem, a 300-iteration cap, and loss of earlier detail when context is compacted at 170k tokens (AgentConn).
Business and practitioner implications
- Cost: a narrow distillation or LoRA now costs tens of dollars in compute. Writing the prompt and doing evaluation are the main costs left.
- Licensing: the Pocket Rewriter and the Qwen-Image 2.1 LoRAs come under the non-commercial Qwen Research License (AI Weekly).
- How to use it: cap spending, require a baseline and a smoke test, and read the model card's evaluation methodology before trusting the headline number.
Read the full section
- Cost: a narrow distillation or LoRA now costs tens of dollars in compute. Writing the prompt and doing evaluation are the main costs left.
- Licensing: the Pocket Rewriter and the Qwen-Image 2.1 LoRAs come under the non-commercial Qwen Research License (AI Weekly). Agate 4-step is MIT-licensed and Citrus Doctor is CC BY 4.0, so check the license before deploying any of them.
- How to use it: cap spending, require a baseline and a smoke test, and read the model card's evaluation methodology before trusting the headline number.