Leviathan:让 Agent 在百万行数据集里只花 436 个 token 找到答案
Leviathan: Agent Deep Memory Over Large Datasets — 436 Tokens at 1M Records
Agent 遇到大型历史数据集时有一个经典困境:直接把所有记录塞进上下文撑爆 token,还是靠 grep 找关键词但不知道找多少才够?
Leviathan 给出了一个不同的答案:本地建索引,Agent 查询时只返回相关的几条记录卡片,token 控制在几百以内,不管数据集多大。
GitHub: https://github.com/elstongun/leviathan | ⭐ 656 | Apache-2.0 | Rust
核心基准(1M 条记录,678MB)
作者提供了一个具体的性能数字,场景是合成的维护日志(一类典型的 Agent 历史数据),数据集规模从 1K 到 1M 条。
| 指标 | Leviathan | 最优 grep 策略 |
|---|---|---|
| 中位 token 消耗 | 436 | 107,122(差 245×) |
| 最差情况(1200 次问答) | 602 tokens | 9,700,000 tokens |
| top-5 相关记录召回率 | 99.0% | 96.0%(30K 字符输出内) |
| rank 1 准确率 | 98.5% | — |
| 中位延迟 | 33ms | 92ms |
需要注意:这是合成数据集上的结果,作者在文档里明确说明这不是标准 benchmark,测试方法见 docs/BENCHMARKS.md,ranking 改动需要附带前后对比数字才能被 merge。
怎么工作
底层是 SQLite + FTS5 全文索引 + BM25 排序。
一次查询的流程:
- 解析 group 参数(精确匹配 → 名称 → 包含 → 模糊)
- FTS5 匹配:group 值和 filter 值作为 token 被索引进去(不是查完再过滤),匹配更精确
- BM25 × boost 排序
- 只解码 top N 条,输出包含
shown N of M的卡片
Agent 拿到的结果类似:
leviathan search · customer C-ACME "Acme Corp" (7 tickets) · 查询 "sso 登录密码重置后跳回登录页"
[1] T-1001 · 2024-01-08 · rel 16.9
密码重置后登录页无限循环
状态: 已关闭 · 优先级: 高
解决方案: 密码重置时清除了旧 session cookie,已发布在 4.2.1。临时方案:清除站点数据。
每张卡片大约 450 tokens,不管数据集是 1K 还是 1M 条。
数据源支持和配置
Leviathan 本身不持有数据库凭据,通过各数据库自己的 CLI 接入:
# CSV / JSONL 直接索引
leviathan index tickets.csv --id "Ticket ID" --group customer_id --date created_at
# PostgreSQL
psql "$DATABASE_URL" -At -c "SELECT row_to_json(t) FROM tickets t" | leviathan index - -c tickets.toml
# SQLite
leviathan index app.db --sql "SELECT * FROM tickets" -c tickets.toml
# DuckDB / Parquet
duckdb -json -c "SELECT * FROM 'events/*.parquet'" | leviathan index - -c events.toml
支持的格式:JSONL、JSON、CSV/TSV、SQLite、任意数据库 CLI 的 JSON 输出。
字段映射(leviathan.toml):
| 字段 | 作用 |
|---|---|
id | 必填,用于 get/upsert/delete 和引用 |
title/text | 卡片标题(权重 2×)/ 被搜索的文本 |
group/group_name | -g 范围搜索和名称解析 |
date | --since/--until 时间过滤 |
filters/display | --where facet 过滤 / 卡片展示字段 |
Agent 接入方式
CLI + Skill(推荐):把 skills/leviathan/SKILL.md 复制到 ~/.claude/skills/ 或粘贴进 AGENTS.md。不用的时候 0 token,用到时再加载。
MCP Server:4 个只读 stdio 工具:
| 工具 | 作用 |
|---|---|
search | 排序全文检索(支持 group/filter/时间范围) |
resolve_group | 解析模糊 group 名称到精确 ID |
get | 通过 ID 获取完整记录 |
describe | 返回数据集摘要(字段/group 列表/示例调用,~640 tokens) |
启动 MCP 服务器并生成对应 Agent 的配置:
leviathan mcp
leviathan wrap claude # 生成 Claude Code 配置
leviathan wrap codex # 生成 Codex 配置
leviathan wrap cursor # 生成 Cursor 配置
一个实际场景:客服工单记忆
Agent 需要回答「某个客户之前报过什么 bug?」这类问题。传统做法:grep 关键词 → 把命中的几百条原始记录全塞进上下文,平均 10 万+ token。
Leviathan 做法:建索引 → 每次查询返回 3–5 张排好序的卡片 → 不超过 600 tokens,99% 的情况下 rank 1 就是答案。
工程细节
- 单静态二进制,2.5MB:
cargo install leviathan-index或下载预编译包,无运行时依赖 - 索引是单个 SQLite 文件:便携,可以 git 追踪或随数据集分发
- 增量更新:
upsert/delete保持索引新鲜,内容未变化时跳过重建 - 原子构建:构建是原子操作,不会出现中间状态
- Exit code 规范:0 ok(零结果也是 0);1 error;2 bad request;3 group 不明确——Agent 可靠判断
- 安全:只读,离线,无
unsafe代码
已知边界
- 这是 v0.1.0,3 天前刚发,中文社区暂无实测记录
- Benchmark 基于单一合成数据集(维护日志)——真实场景的泛化能力需要自己验证
- FTS5 是词项匹配,语义相似(embedding 检索)不在这个工具的覆盖范围
Apache-2.0 开源。elstongun 维护,Rust 单二进制,3 天 656 stars。开源仅供学习参考。
Leviathan: Agent Deep Memory Over Large Datasets — 436 Tokens at 1M Records
Agents working over large historical datasets face a classic dilemma: stuff everything into context (expensive) or grep for keywords (noisy and token-heavy). Leviathan takes a different approach: build a local full-text index once, and return only ranked result cards at query time — a few hundred tokens regardless of dataset size.
GitHub: https://github.com/elstongun/leviathan | ⭐ 656 | Apache-2.0 | Rust
Core Benchmark (1M records, 678MB)
Tested on a synthetic maintenance log. Dataset scaled from 1K to 1M records.
| Metric | Leviathan | Best grep strategy |
|---|---|---|
| Median tokens/question | 436 | 107,122 (245× more) |
| Worst case (1,200 questions) | 602 tokens | 9,700,000 tokens |
| Top-5 relevant record return | 99.0% | 96.0% (within 30K chars) |
| Rank 1 accuracy | 98.5% | — |
| Median latency | 33ms | 92ms |
Caveat: single synthetic dataset; methodology in docs/BENCHMARKS.md; ranking PRs require before/after benchmark numbers.
How It Works
SQLite + FTS5 + BM25 × boost ranking. A query: resolve group → FTS5 match (group/filter values are indexed tokens, not post-filters) → rank → decode top N into capped cards (~450 tokens each). Every response includes shown N of M so agents can distinguish “no match” from “no data.”
Data Sources
No credentials held by Leviathan — pipe from any database CLI:
leviathan index tickets.csv --id "Ticket ID" --group customer_id --date created_at
psql "$DATABASE_URL" -At -c "SELECT row_to_json(t) FROM tickets t" | leviathan index -
leviathan index app.db --sql "SELECT * FROM tickets" -c tickets.toml
duckdb -json -c "SELECT * FROM 'events/*.parquet'" | leviathan index -
Agent Integration
Recommended: CLI + SKILL.md — copy skills/leviathan/SKILL.md into ~/.claude/skills/. Zero tokens until invoked.
Optional: MCP server — 4 read-only stdio tools (search, resolve_group, get, describe). leviathan wrap claude|codex|cursor|... generates agent config.
Engineering Notes
- Single static binary, 2.5MB, no runtime dependencies
- Index is a single portable SQLite file
- Incremental updates via
upsert/delete; skips rebuild when nothing changed - Atomic index builds; no intermediate states
- Deterministic exit codes (0/1/2/3) for reliable agent parsing
- Read-only, offline, zero
unsafecode
Boundaries
- v0.1.0 released 2026-10-06; no real-world validation beyond the synthetic benchmark yet
- FTS5 is term-based; semantic/embedding search is out of scope
- Group resolution fails loudly (exit 3) rather than silently guessing — correct for agent use, requires explicit handling
Apache-2.0. Maintained by elstongun. Single Rust binary, 656 stars in 3 days. For technical reference only.
关于本站 · 免责声明
🍄 Mushroom Research Blog 是非营利、免费公开的个人科技观察博客与公众号 XStack18,不接受商业合作、不代表任何企业或机构立场,也不谋求商业利益。我们以个人视角客观中立地记录和分析 AI、Web3 等领域的最新模型发布与技术动态——不止转述新闻标题或二手信息,而是给出有独立思考的深入分析,希望帮更多人获得有价值的一手科技认知。
⚠️ 文中介绍的开源代码与模型,仅供学习交流与技术借鉴。它们大多仍处于早期阶段,有待进一步研究和验证,请勿直接用于工作或生产环境;如需采用,请先自行充分测试,并核实其许可证与安全性。
Open-source code and models featured here are shared for learning and reference only. Most are early-stage and still need further study and verification — please don't use them directly in your work or in production. Test them thoroughly and check their licenses and security first.
- 本站文章均为作者基于公开信息的个人研究与观点整理,不代表文中提及的任何公司、产品、模型的官方立场,未与其构成商业关联或合作关系。
- 科技行业信息更新极快,我们尽力保证内容准确、及时,但不对完整性、实时性做绝对保证,具体请以相关企业/项目官方公告为准。
- 文中引用的第三方商标、产品名称、图片、数据等版权归原权利人所有,我们会尽量注明来源;如你认为存在版权疑问或侵权,请通过下方邮箱联系我们,收到通知后会尽快核实处理(更正、加注来源或删除)。
- 文章内容仅为技术科普与个人观点,不构成投资、法律或其他专业建议,据此进行任何决策的后果需自行判断和承担。
📮 侵权 / 勘误 / 合作咨询:[email protected]
💬 评论与讨论
使用 GitHub 账号登录后发表评论