What's New
Eval Platform (7 subcommands)
Full-featured evaluation system beyond basic lm-evaluation-harness wrapper:
soup eval benchmark— standard benchmarks (mmlu, gsm8k, hellaswag) via lm-evaluation-harnesssoup eval custom --tasks eval.jsonl— custom eval tasks with 4 scoring modes (exact, contains, regex, semantic)soup eval judge— LLM-as-a-judge scoring (OpenAI, Ollama, local server backends)soup eval auto— automatic evaluation after training viaevalconfig in soup.yamlsoup eval compare <run1> <run2>— side-by-side eval comparison with regression detectionsoup eval leaderboard— local model leaderboard with JSON/CSV exportsoup eval human— terminal A/B comparison with Elo ratings
Auto-Eval Config
eval:
auto_eval: true
benchmarks: [mmlu, gsm8k]
custom_tasks: eval_tasks.jsonlSecurity
- Custom eval JSONL: schema validation, capped at 10k tasks
- Regex scoring: ReDoS guard (pattern + input length caps)
- Judge API: SSRF protection (localhost-only HTTP), API key isolation per provider
- Human eval: local-only terminal UI, 10k prompt cap
- Leaderboard: read-only SQLite queries
Stats
- 1585 tests across 58 test files
- 60.6% coverage
- ruff lint clean
Install / Upgrade
pip install --upgrade soup-cli