分层记忆架构

📎 引用文件

本文引用的文件 - persistent.py - hierarchy.py - lifecycle.py - compression.py - search_index.py - semantic_links.py - test_memory_lifecycle.py - test_persistent_memory.py

目录

  1. 简介
  2. 项目结构
  3. 核心组件
  4. 架构总览
  5. 详细组件分析
  6. 依赖关系分析
  7. 性能考量
  8. 故障排查指南
  9. 结论
  10. 附录:使用示例与最佳实践

简介

本文件系统性阐述 Vibe-Trading 的分层记忆系统设计理念与实现机制。该系统以“持久化存储 + 层次化路由 + 生命周期管理 + 压缩归档 + 全文检索 + 语义关联”为核心,提供跨会话的记忆能力,支持在用户、反馈、项目、参考等类别之间组织记忆,并通过质量评分、访问衰减、垃圾回收与自动压缩,实现容量控制与长期可维护性。同时,通过 SQLite FTS5 全文索引与 BM25 语义链接,提升查询效率与上下文召回质量。

项目结构

记忆子系统位于 agent/src/memory,主要模块如下: - persistent.py:持久化存储、条目模型、基础读写、去重、索引快照、重要性计算 - hierarchy.py:按 memory_type 分类的目录路由与扫描优化 - lifecycle.py:质量强化、访问追踪、重要性衰减、垃圾回收(归档/删除) - compression.py:三级压缩(raw → daily → digest),基于 TF-IDF 关键句提取与摘要生成 - search_index.py:SQLite FTS5 全文检索索引,CJK 分词与查询增强 - semantic_links.py:BM25 语义链接发现与 .relations.json 侧边文件管理

graph TB PM["PersistentMemory<br/>持久化存储"] --> H["MemoryHierarchy<br/>层次化路由"] PM --> SI["MemorySearchIndex<br/>FTS5 全文索引"] PM --> SL["SemanticLinker<br/>语义链接"] LC["MemoryLifecycle<br/>生命周期管理"] --> PM LC --> CP["CompressionPipeline<br/>压缩流水线"] PM --> |写入/读取| FS["文件系统<br/>.md 条目 + 索引"] SI --> DB["SQLite 数据库<br/>memory_index.db"] SL --> FS

图表来源 - persistent.py:196-637 - hierarchy.py:34-436 - lifecycle.py:71-421 - compression.py:160-353 - search_index.py:113-481 - semantic_links.py:158-372

章节来源 - persistent.py:196-637 - hierarchy.py:34-436 - lifecycle.py:71-421 - compression.py:160-353 - search_index.py:113-481 - semantic_links.py:158-372

核心组件

章节来源 - persistent.py:122-637 - hierarchy.py:34-436 - lifecycle.py:71-421 - compression.py:160-353 - search_index.py:113-481 - semantic_links.py:158-372

架构总览

分层记忆系统围绕“写路径”和“读路径”展开: - 写路径:add() 生成 frontmatter 与 body,必要时通过 hierarchy.route_entry() 路由到类别目录;更新 MEMORY.md 索引;可选写入 FTS5 索引与语义链接;并发写通过文件锁保护。 - 读路径:list_entries() 扫描所有 .md(支持层级扫描),解析 frontmatter 并计算 importance;find_relevant() 优先走 FTS5 索引,否则回退到 token 加权匹配;可选扩展语义链接结果。 - 生命周期:reinforce() 调整 quality_score;track_access() 更新访问计数;run_gc() 依据阈值归档或删除低重要性条目,并在非 dry_run 时触发压缩流水线。 - 压缩:根据 last_accessed 与当前时间差决定目标级别(daily/digest),先归档原文再重写 frontmatter 与 body。

sequenceDiagram participant App as "调用方" participant PM as "PersistentMemory" participant H as "MemoryHierarchy" participant SI as "MemorySearchIndex" participant SL as "SemanticLinker" participant FS as "文件系统" App->>PM : add(name, content, type, description) PM->>H : route_entry(type, slug.md) H-->>PM : path PM->>FS : 写入 frontmatter + body PM->>SI : index_entry(id, title, desc, keywords, body) PM->>SL : discover_links(...) SL-->>PM : links PM->>FS : 保存 .relations.json PM-->>App : path

图表来源 - persistent.py:462-578 - hierarchy.py:70-90 - search_index.py:207-241 - semantic_links.py:179-230

章节来源 - persistent.py:462-578 - hierarchy.py:70-90 - search_index.py:207-241 - semantic_links.py:179-230

详细组件分析

持久化存储(PersistentMemory)

flowchart TD Start(["写入入口 add"]) --> CheckDup["检查重复(30s窗口)"] CheckDup --> |重复| ReturnNone["返回 None"] CheckDup --> |不重复| Route["层级路由 route_entry"] Route --> WriteFM["写入 frontmatter + body"] WriteFM --> UpdateIndex["更新 MEMORY.md 索引"] UpdateIndex --> OptionalLinks{"是否启用语义链接?"} OptionalLinks --> |是| Discover["发现并保存 .relations.json"] OptionalLinks --> |否| Done["完成"] Discover --> Done

图表来源 - persistent.py:440-578 - hierarchy.py:70-90

章节来源 - persistent.py:122-637

层次化路由(MemoryHierarchy)

classDiagram class MemoryHierarchy { +base_dir Path +route_entry(memory_type, filename) Path +recover_extensionless_entries() Path[] +scan_all() Path[] +scan_category(category) Path[] +rebuild_index(entries) void +prune_search_scope(query_tokens, category_filter) Path[] +migrate_flat_entry(file_path, memory_type) Path? }

图表来源 - hierarchy.py:34-436

章节来源 - hierarchy.py:34-436

生命周期管理(MemoryLifecycle)

flowchart TD GCStart["运行垃圾回收"] --> Scan["列出所有条目"] Scan --> ForEach{"遍历条目"} ForEach --> AgeCheck{"年龄 >= MIN_AGE_DAYS?"} AgeCheck --> |否| Next["下一个条目"] AgeCheck --> |是| ImpCalc["计算重要性"] ImpCalc --> Threshold{"低于归档/删除阈值?"} Threshold --> |是| Action{"归档 or 删除"} Action --> DryRun{"dry_run?"} DryRun --> |是| Log["记录日志"] DryRun --> |否| Exec["执行动作"] Exec --> Compress{"是否启用压缩?"} Compress --> |是| Trigger["触发压缩流水线"] Compress --> |否| Next Log --> Next Next --> ForEach

图表来源 - lifecycle.py:183-273 - persistent.py:80-91

章节来源 - lifecycle.py:71-421 - persistent.py:80-91

压缩流水线(CompressionPipeline)

flowchart TD Entry["记忆条目"] --> Decide{"should_compress(level, last_accessed, now)"} Decide --> |raw -> daily| Archive1["归档原文"] Decide --> |daily -> digest| Archive2["归档原文"] Archive1 --> CompressDaily["compress_to_daily()"] Archive2 --> CompressDigest["compress_to_digest()"] CompressDaily --> Retention["estimate_retention()"] CompressDigest --> Retention Retention --> WriteBack["_write_compressed() 更新 frontmatter + body"]

图表来源 - compression.py:168-353 - lifecycle.py:323-379

章节来源 - compression.py:160-353 - lifecycle.py:323-379

全文检索(MemorySearchIndex)

sequenceDiagram participant PM as "PersistentMemory" participant SI as "MemorySearchIndex" participant DB as "SQLite" PM->>SI : index_entry(id, title, desc, keywords, body) SI->>DB : INSERT OR REPLACE (memories) Note over SI,DB : FTS5 触发器自动同步 memories_fts PM->>SI : search(query) SI->>DB : MATCH query with sanitized tokens DB-->>SI : ranked results SI-->>PM : MemoryMatch[]

图表来源 - search_index.py:147-196 - search_index.py:207-241 - search_index.py:252-302

章节来源 - search_index.py:113-481

语义链接(SemanticLinker)

classDiagram class SemanticLinker { +discover_links(entry_title, entry_tokens, all_entries_data, top_k) Tuple[] +save_relations(entry_path, links) void +load_relations(entry_path) Tuple[] +resolve_wikilinks(body) str[] +get_relation_path(entry_path) Path +remove_relations(entry_path) void }

图表来源 - semantic_links.py:158-372

章节来源 - semantic_links.py:158-372

依赖关系分析

graph LR PM["PersistentMemory"] --> H["MemoryHierarchy"] PM --> SI["MemorySearchIndex"] PM --> SL["SemanticLinker"] LC["MemoryLifecycle"] --> PM LC --> CP["CompressionPipeline"] SI --> DB["SQLite"] SL --> FS["文件系统"] PM --> FS

图表来源 - persistent.py:196-637 - lifecycle.py:71-421 - search_index.py:113-481 - semantic_links.py:158-372

章节来源 - persistent.py:196-637 - lifecycle.py:71-421 - search_index.py:113-481 - semantic_links.py:158-372

性能考量

[本节为通用性能讨论,无需特定文件来源]

故障排查指南

章节来源 - test_memory_lifecycle.py:228-305 - test_persistent_memory.py:285-340 - search_index.py:252-302 - semantic_links.py:321-344

结论

Vibe-Trading 的分层记忆系统通过清晰的职责划分与模块化设计,实现了高可用、可扩展、可维护的记忆能力。层级路由与全文检索显著提升查询效率;生命周期管理与压缩流水线保障长期存储健康;语义链接增强上下文召回质量。整体方案在保证一致性与并发安全的前提下,提供了灵活的配置开关与完善的错误处理,适用于复杂交易研究场景中的知识沉淀与复用。

[本节为总结性内容,无需特定文件来源]

附录:使用示例与最佳实践

章节来源 - test_persistent_memory.py:124-245 - test_memory_lifecycle.py:228-305 - persistent.py:462-578 - lifecycle.py:183-273 - compression.py:168-353