What's New
Embedding Model Fine-Tuning
task: embedding— fine-tune sentence embedding models (BGE, E5, GTE, INSTRUCTOR)- Contrastive loss (InfoNCE with configurable temperature)
- Triplet loss with configurable margin
- Cosine loss for simple similarity training
- Configurable pooling:
mean,cls,lasttoken - New template:
soup init --template embedding
ONNX Export
soup export --format onnx— export models via optimum- Supports
--onnx-taskfor text-generation or feature-extraction - Install:
pip install 'soup-cli[onnx]'
TensorRT-LLM Export
soup export --format tensorrt— export for high-throughput GPU inference- Two-step pipeline: HF checkpoint → TensorRT engine
- Install:
pip install 'soup-cli[tensorrt]'
Speculative Decoding
soup serve --speculative-decoding <draft-model>— 2-3x faster generation- Transformers backend: HuggingFace assisted generation
- vLLM backend: native speculative decoding
- SSRF protection on draft model paths
Security
- ONNX export: no unconditional
trust_remote_code - Speculative decoding: URL validation, warning panel before draft model load
- TensorRT export: separate error handling per subprocess call
- Embedding config: Literal constraints, margin validation
Stats
- 1270 tests across 52 test files
- 58% code coverage
- All ruff lint rules pass
Install / Upgrade
pip install --upgrade soup-cli