github MakazhanAlpamys/Soup v0.19.0
v0.19.0 — Eval Platform

latest releases: v0.75.2, v0.75.1, v0.75.0...
6 months ago

What's New

Eval Platform (7 subcommands)

Full-featured evaluation system beyond basic lm-evaluation-harness wrapper:

  • soup eval benchmark — standard benchmarks (mmlu, gsm8k, hellaswag) via lm-evaluation-harness
  • soup eval custom --tasks eval.jsonl — custom eval tasks with 4 scoring modes (exact, contains, regex, semantic)
  • soup eval judge — LLM-as-a-judge scoring (OpenAI, Ollama, local server backends)
  • soup eval auto — automatic evaluation after training via eval config in soup.yaml
  • soup eval compare <run1> <run2> — side-by-side eval comparison with regression detection
  • soup eval leaderboard — local model leaderboard with JSON/CSV export
  • soup eval human — terminal A/B comparison with Elo ratings

Auto-Eval Config

eval:
  auto_eval: true
  benchmarks: [mmlu, gsm8k]
  custom_tasks: eval_tasks.jsonl

Security

  • Custom eval JSONL: schema validation, capped at 10k tasks
  • Regex scoring: ReDoS guard (pattern + input length caps)
  • Judge API: SSRF protection (localhost-only HTTP), API key isolation per provider
  • Human eval: local-only terminal UI, 10k prompt cap
  • Leaderboard: read-only SQLite queries

Stats

  • 1585 tests across 58 test files
  • 60.6% coverage
  • ruff lint clean

Install / Upgrade

pip install --upgrade soup-cli

Don't miss a new Soup release

NewReleases is sending notifications on new releases.