v1.6.2 — Embedding reliability & Slack history
Release Date: 2026-09-29
Changes: v1.6.1 → v1.6.2
Pull Requests: #4932, #4951, #5146, #5196, #5231, #5235, #5236, #5240, #5241, #5242, #5247, #5250, #5251, #5274
Summary
This release focuses on making embeddings reliable and transparent (no more silent truncation), adds Slack conversation import and sync, and reorganizes how 'skills' are stored for better cross-platform tooling. You'll see clearer errors when embedding fails, safer chunk sizing based on model token limits, and updated CLI/docs to match these behaviors.
Highlights
- Embedding fixes prevent silent truncation and size chunks by each model's token limit so your content is embedded intact and consistently.
- Slack integration: import and continuously sync conversation history from Slack into Cognee.
- Skills moved to a new .agents/skills location with Windows linking support and recovery when linking fails.
- Better errors and faster failure on unrecoverable embedding errors — you’ll see real causes instead of opaque messages.
- CLI and docs refreshed (including database migration/commands and Docker quickstart auth fixes).
Breaking Changes
- Skills path change: Built-in and packaged skills are now under .agents/skills (moved from earlier locations). If you maintain custom scripts, automation, build hooks, or CI that reference the old paths, update them to .agents/skills. A small helper script (link-skills.sh) and a Windows hook were added to assist linking.
- env_file removed from settings: The runtime setting that automatically loaded an env_file was removed. If you relied on that behavior to load environment variables via the settings file, switch to an explicit env loader in your start scripts or use the newly added shared/env_file helper (see docs/tests for examples).
New Features
- Slack conversation import and sync — A new integration to import and keep Slack conversation history in sync with Cognee. This pulls channel and thread messages into your datasets so search and memory features can use Slack content. Why it matters: you can build memories, search, and summaries from historical Slack data without manual export.
- Chunks sized by model input limit — The system now measures each embedding model's input limit (the maximum number of tokens a model accepts) and splits text into chunks sized to that limit. What it does: chunking no longer relies on a fixed heuristic; it respects the actual model/tokenizer limits to avoid unexpected truncation. Why it matters: search and memory will be more complete and consistent because documents are not silently shortened when sent to an embedding model.
- Skills directory reorganization (.agents) and Windows link hook — Skills (automated agent scripts and examples) are moved from older locations into a new .agents/skills directory and a Windows-friendly hook was added for linking. What it does: standardizes where built-in skills live and makes the linking process work on Windows. Why it matters: easier, more consistent workflows for installing, linking, and running skills across platforms.
Improvements
- Robust embedding input limit resolution — The SDK now resolves a model's input limit even when optional tokenizer libraries (like transformers) are not installed, and does so asynchronously where appropriate. This reduces startup surprises and improves behavior for lightweight installs.
- Accurate token counting across tokenizers — Token counts now use each provider's tokenizers (including HuggingFace tokenizers when available), and the model's own special tokens are included in counts. What it does: improves chunk sizing and prevents off-by-a-few-token truncation.
- Fastembed batch padding — Batches sent to the fast embedding backend are now padded to the longest text in the batch to avoid inconsistent batch-level truncation. Result: fewer silent truncations and more predictable embeddings when sending mixed-length inputs.
- Clearer embedding errors and fail-fast behavior — Embedding failures now show the underlying cause (not just a generic message) and terminal errors stop retries quickly. Why it matters: you can debug and recover faster when an external provider or configuration is wrong.
- Recover skill links on failure and validate links without keys — The linking process for skills is more robust: partial failures try to recover links and the link-checking workflow can run without a secret key. This reduces brittle operations and makes CI checks simpler.
- CLI docs and behavior updates — The CLI documentation has been expanded and clarified (database migration commands, new examples, memory commands notes). The Docker quickstart 401 trap has been improved so the quickstart behaves correctly when auth is missing.
Performance
- More reliable embedding throughput — By padding fastembed batches and sizing chunks to the model input limits, embeddings avoid silent truncation that used to reduce effective throughput/quality. This improves the practical throughput of usable embeddings for mixed-length inputs.
- Faster detection of embedding limits in lightweight setups — Input limit resolution now works without heavyweight tokenizer dependencies and resolves asynchronously, reducing delays during embedding initialization.
Security
- Docker quickstart 401 handling — The Docker quickstart and related docs were fixed to avoid a confusing 401 trap and now produce clearer authentication guidance.
Bug Fixes
- Marked optional transformers imports correctly for type checking so optional dependencies behave cleanly during installs.
- Counted HuggingFace tokens correctly using tokenizers (avoids miscounting when different tokenizer code paths are used).
- Read a model's input limit even when transformers is not installed, preventing silent fallbacks that caused wrong chunk sizes.
- Ignored previously-saved truncation/padding metadata that could corrupt future chunk sizing decisions.
- Recovered skill links when clearing fails (link recovery to avoid broken skill installations).
- Dropped an env_file setting from runtime settings to avoid surprising behavior (see Breaking changes).
- Logged full tracebacks on delegated CLI command failures so errors are easier to diagnose.
- Changed the order in which databases are pruned to fix an edge-case deletion ordering bug.
Technical Changes
- Refactored embedding input_limit() into the engine interface and kept each limit lookup with its provider to make limit resolution pluggable.
- Dropped a no-op re-raise in fastembed and simplified retry predicate classification to handle missing embedding SDKs cleanly.
- Refactored tokenizer ordering guards (removed transformers-before-torch guard) and added guarded tests to prevent regressions in import order and padding behavior.
- Large suite of tests and example additions: many end-to-end, unit, and journey tests were added or expanded (embedding limits, slack history, chunking, journeys), plus a new company_brain multi-source demo and expanded CI workflows.
- Documentation and skills: numerous documentation updates (AGENTS.md, skills docs, migration guidance, README alignment) and formatting fixes for Python skill examples.
Dependency Updates
Updated:
- enola-cli: ==0.4.21 → ==0.4.26
Compatibility
| Component | Supported / Required |
|---|---|
| Python | >=3.10,<3.15
|
| pydantic | >=2.10.5
|
| litellm | >=1.83.7,<1.97.0
|
| fastapi | >=0.116.2,<1.0.0
|
| sqlalchemy | >=2.0.39,<3.0.0
|
| lancedb | >=0.24.3,<1.0.0
|
| ladybug | ==0.19.0
|
— The Cognee Team · 2026-09-29