人工智能AI 应用大模型RAG后端AI Agent【免费下载链接】khojYour AI second brain. Self-hostable. Get answers from the web or your docs. Build custom agents, schedule automations, do deep research. Turn any online or local LLM into your personal, autonomous AI (gpt, claude, gemini, llama, qwen, mistral). Get started - free.项目地址https://gitcode.com/GitHub_Trending/kh/khoj点击查看免费下载导读本文以 Khoj 仓库测试数据集中的一篇典型个人笔记tests/data/markdown/Meet Arun and Pablo for Lunch.markdown为贯穿线索完整拆解 Khoj 将 Markdown 笔记转化为可被 AI 搜索与聊天引用的语义条目的处理链路从文件解析、标题拆分、条目构造到向量化入库与日期索引再到用自然语言 dt日期过滤器回查。读完本文你将掌握 Khoj 的 Markdown 内容管线MarkdownToEntries的核心机制并能在自托管部署中预判笔记如何被切分、索引与检索。一、样本文件剖析个人笔记中常见的三类信息载体Khoj 定位为 Your AI second brain其核心能力之一就是把散落在本地磁盘、桌面同步目录中的个人 Markdown 文件变成可供语义搜索和聊天引用的知识库。仓库测试数据中收录了一批模拟真实个人笔记的文件其中Meet Arun and Pablo for Lunch.markdown极具代表性全文如下--- SCHEDULED: 2023-04-01 CLOSED: 2023-04-01 --- Met Pablo and Arun for Lunch at Arak, Medellin. Arun just sold his apartment in Nairobi and is moving with his wife to Medellin in April 2023! Pablo mentioned his son Amal just got admission into the Colegio Superior de Gastronomia in Mexico City. Last of his 3 kids to leave the nest! 2023-04-01 Arak Dosa for Lunch Expenses:Food:Dining 11.00 USD这个文件浓缩了个人知识管理场景中三类高频信息载体YAML/Org 风格 frontmatter 元数据SCHEDULED/CLOSED字段记录了事件的计划与完成日期是 Khoj 日期索引的天然素材自由文本正文自然语言记录的见闻与事实是语义检索的主要对象账本格式交易行Ledger/Beancount 风格2023-04-01 Arak Dosa for Lunch与Expenses:Food:Dining 11.00 USD同样携带可被解析的结构化日期。Khoj 会逐行处理这些内容而不同类型的字段会走不同的子管线。下面依次展开。二、入口与主流程MarkdownToEntries 的处理骨架在 Khoj 的源码中Markdown 文件接入的统一入口是MarkdownToEntries.process()位于 markdown_to_entries.py。该方法接收一个{文件路径: 文件内容}的字典依次完成三个阶段def process(self, files: dict[str, str], user: KhojUser, regenerate: bool False) - Tuple[int, int]: deletion_file_names set([file for file in files if files[file] ]) files_to_process set(files) - deletion_file_names ... max_tokens 256 with timer(Extract entries from specified Markdown files, logger): file_to_text_map, current_entries MarkdownToEntries.extract_markdown_entries(files, max_tokens) with timer(Split entries by max token size supported by model, logger): current_entries self.split_entries_by_max_tokens(current_entries, max_tokens) with timer(Identify new or updated entries, logger): num_new_embeddings, num_deleted_embeddings self.update_embeddings(...)三个阶段的职责分别是阶段方法作用条目抽取extract_markdown_entries按标题结构把单个 Markdown 文件切成一个或多个Entryraw 原文 行号 URI超长拆分split_entries_by_max_tokens对超过模型 token 上限的 compiled 文本做递归切块保证每条不超过 256 tokens增量入库update_embeddings计算 MD5 哈希、生成向量嵌入、增量写入数据库、索引日期一个值得注意的细节process中默认max_tokens 256而测试用例test_markdown_to_entries.py在直接调用extract_markdown_entries时常用max_tokens3强制触发递归拆分以便验证拆分边界行为。也就是说256 是生产默认值而 3/10/12 是测试专用的小值目的是把大文件拆小的逻辑暴露出来。2.1 空文件即删除信号入口处还有一个易被忽略的约定deletion_file_names set([file for file in files if files[file] ])——内容为空的文件路径会被解释为删除该文件的索引不会进入抽取流程而是最终在update_embeddings阶段通过EntryAdapters.delete_entry_by_file删除对应条目。这意味着客户端删除文件时只需向服务端上报一个空内容条目即可触发索引清理。三、条目切分原理按标题层级递归、保留祖先标题extract_markdown_entries遍历每个文件调用process_single_markdown_filemarkdown_to_entries.py完成真正的拆分。其核心算法如下先拼接标题祖先把当前 section 的所有祖先标题按层级顺序用#前缀拼在正文之前ancestry_string保证每个条目自带上文语境判定是否直接成条若拼接后的内容 token 数 ≤max_tokens且正文中不存在比当前层级更深的子标题则整个内容作为一个条目否则递归拆分按下一个存在的标题层级用正则re.split(rf(\n|^)(?[#]{{{next_heading_level}}} .\n?), ...)切出多个 section对每个 section 更新其标题祖先current_ancestry[next_heading_level] current_section_title再递归处理同时用current_line_offset精确追踪每个 section 在文件中的起始行号。对应的行为在测试中有非常明确的断言test_markdown_to_entries.pydef test_extract_entries_with_non_incremental_heading_levels(tmp_path): ... assert entries[1][0].raw # Heading 1\n#### Sub-Heading 1.1, Ensure entry includes heading ancestory assert entries[1][1].raw # Heading 1\n## Sub-Heading 1.2, Ensure entry includes heading ancestory即子标题条目会完整继承其所有祖先标题即使标题层级跳跃#直接跳到####也能正确处理。3.1 无标题文件的处理恰适用于我们的样本文件Meet Arun and Pablo for Lunch.markdown全文没有任何#标题因此它的处理路径对应测试 test_extract_markdown_with_no_headingsdef test_extract_markdown_with_no_headings(tmp_path): ... # Ensure raw entry with no headings do not get heading prefix prepended assert not entries[1][0].raw.startswith(#) # Ensure compiled entry has filename prepended as top level heading assert entries[1][0].compiled.startswith(expected_heading)可以推断这个样本文件被抽取后会形成一个单一条目raw保持原文不加标题前缀而compiled真正喂给嵌入模型和 LLM 的字段会被自动加上文件名作为一级标题。这一点在convert_markdown_entries_to_maps中有直接实现markdown_to_entries.pyheading parsed_entry.splitlines()[0] if re.search(r^#\s, parsed_entry) else # Append base filename to compiled entry for context to model prefix f# {entry_filename}\n# if heading else f# {entry_filename}\n compiled_entry f{prefix}{parsed_entry}于是入库后的compiled大致形如# Meet Arun and Pablo for Lunch.markdown Met Pablo and Arun for Lunch at Arak, Medellin. ...文件名作为顶级标题这一设计有实际价值当 AI 引用该条目时模型能直接看到内容来源文件名回答中可自然带出来自某篇笔记的上下文。3.2 行号溯源每个条目携带精确文件定位在 convert_markdown_entries_to_maps 中每个条目还会生成形如file://{绝对路径}#lineN的 URI其中N是条目在源文件中的起始行号1 基。本地文件走file://前缀URL 则原样保存。test_line_number_tracking_in_recursive_split 专门验证了这一点它用max_tokens10强制对大文件做多层递归拆分然后逐个回读原始文件校验每个条目 URI 中的行号所指向的行内容与条目首行一致。这意味着无论文件被拆得多碎用户都能从检索结果跳回到笔记中的准确位置——这是第二大脑类产品体验的关键细节。四、超长条目的二次拆分与清洗当单个条目的 compiled 文本超过模型 token 上限时split_entries_by_max_tokenstext_to_entries.py会兜底处理。它使用langchain_text_splitters.RecursiveCharacterTextSplitter切分优先级为text_splitter RecursiveCharacterTextSplitter( chunk_sizemax_tokens, separators[\n\n, \n, !, ?, ., , \t, ], keep_separatorTrue, length_functionlambda chunk: len(TextToEntries.tokenizer(chunk)), chunk_overlap0, )切块顺序是段落 行 感叹号 问号 句号 空格 制表符 字符tokenizer即简单的text.split()按空白切词。额外的细节包括从原始raw文本中反查每个 chunk 的实际位置重算#line行号保证切块后的 URI 依然精确text_to_entries.py除首个 chunk 外后续 chunk 会前置截短的条目标题取标题最后 100 字符text_to_entries.py让模型知道这些碎片属于同一主题用remove_long_words丢弃超过 500 字符的超长单词、用clean_field清除\0等非法字符避免污染嵌入质量。五、日期信息如何被提取与利用样本文件中的SCHEDULED: 2023-04-01、CLOSED: 2023-04-01以及账本行首的2023-04-01都属于 Khoj 的日期索引体系。日期提取由DateFilter承担date_filter.py它维护了 20 种正则覆盖结构化日期与自然语言日期结构化\b\d{4}[-\/]\d{2}[-\/]\d{2}\b如2023-04-01、\d{2}[-\/]\d{2}[-\/]\d{4}如01-04-1984等自然语言1st April 1984、April 2021、Apr 84等。对应的行为有测试直接验证test_date_filter.pyextracted_dates DateFilter().extract_dates(head CREATED: today SCHEDULED: 1984-04-01 tail) assert extracted_dates [datetime(1984, 4, 1, 0, 0, 0)], Expected only Y-m-d structured date to be extracted可见SCHEDULED: 2023-04-01这种写法会被稳定地解析出2023-04-01这个日期相对日期如today会被忽略。抽取到的日期在update_embeddings尾部被批量写入EntryDates表text_to_entries.py与条目建立关联构成按时间检索的倒排索引。5.1 查询侧用 dt 过滤器按时间回查入库的日期在查询侧通过dt过滤器消费。DateFilter定义了查询语法date_filter.pydt([:]{1,2})日期表达式常见用法包括查询示例语义dt:2023-04-01命中 2023-04-01 当天dtyesterday dttomorrow昨天 00:00 至明天 00:00 之间dt:2 years ago两年前的当天dt:last week上一自然周比较符、、、、、:会被组合成交集区间date_filter.py。extract_date_range的测试给出了精确语义例如dt:1984-01-01返回[当天0点, 次日0点)test_date_filter.py。因此对本文的样本笔记若你自托管 Khoj 并已同步该文件搜索dt:2023-04-01 午餐即可精准命中2023 年 4 月 1 日那次与 Arun、Pablo 的午餐相关条目。更完整的过滤器说明可参考 query-filters.md 与 search.md。六、增量更新与向量入库哈希驱动的幂等同步update_embeddingstext_to_entries.py是管线的收尾环节其增量策略非常清晰对每条compiled文本计算MD5 哈希hash_func见 text_to_entries.py从数据库按user hashed_value file_type查出已有哈希只对新增哈希生成嵌入embeddings_model[model.name].embed_documents(...)将新条目批量写入Entry表字段包括raw、compiled、heading截断到 1000 字符、file_path、hashed_value、corpus_id、url即file://...#lineN与search_model反过来删除数据库中已不存在于当前文件内容的旧哈希条目to_delete_entry_hashes existing - current同步更新FileObject中保存的文件原文供后续全文检索使用。这一设计意味着重复同步同一文件是幂等的——内容未变则哈希不变、不产生新嵌入内容局部修改则只有受影响条目被重算。process方法还会把新增/删除的条目数返回给调用方markdown_to_entries.py方便上层记录同步状态。值得注意MarkdownToEntries并非只服务本地文件GitHub 数据源处理器也复用了它的process_single_markdown_file与convert_markdown_entries_to_maps见 github_to_entries.py说明按标题拆分与条目构造是 Khoj 所有 Markdown 类内容源的公共底座。七、验证与测试如何用仓库证据复现这条链路如果你想亲手验证以上机制仓库内已备好全部材料样本输入tests/data/markdown/ 目录下 20 个真实风格的笔记文件除本文的主角外还有带SCHEDULED/CLOSED的 Preparing to File Taxes for 2022.markdown、纯账本格式的 Miscellaneous Transactions.markdown、以及用于行号追踪验证的超长文件 main_readme.md单元测试test_markdown_to_entries.py 覆盖无标题、单条目、多条目、不同层级标题、非递增标题层级、标题前文本、小文件单条、递归拆分行号追踪共 8 类场景日期测试test_date_filter.py 覆盖SCHEDULED:解析、dt比较符区间、自然语言日期解析等运行方式仓库使用 pytest 组织测试pytest.ini可执行pytest tests/test_markdown_to_entries.py tests/test_date_filter.py运行上述用例。八、小结一条笔记的完整旅程以Meet Arun and Pablo for Lunch.markdown为样本Khoj 的完整处理链路可归纳为抽取process_single_markdown_file判断文件无标题且体量小 256 tokens整个文件作为单一条目raw保留原文构造convert_markdown_entries_to_maps以文件名为一级标题生成compiled并生成file://...#line1的定位 URI日期索引DateFilter.extract_dates从SCHEDULED/CLOSED/账本行中抽出2023-04-01写入EntryDates向量化update_embeddings对 compiled 文本计算 MD5 与语义向量增量写入Entry表检索查询时用自然语言向量召回 dt日期过滤器做时间限定最终由 AI 引用对应条目作答。理解了这条链路你便能在自托管 Khoj 时预测笔记文件的索引行为正文是否有标题决定条目切分粒度frontmatter 中的结构化日期决定时间检索能力文件是否超长决定是否二次切块。若希望让某类笔记获得更细粒度的检索按小节命中只需在笔记中合理使用##标题即可——这背后正是 markdown_to_entries.py 的标题递归拆分逻辑在起作用。赞分享人工智能AI 应用大模型RAG后端AI Agent【免费下载链接】khojYour AI second brain. Self-hostable. Get answers from the web or your docs. Build custom agents, schedule automations, do deep research. Turn any online or local LLM into your personal, autonomous AI (gpt, claude, gemini, llama, qwen, mistral). Get started - free.项目地址https://gitcode.com/GitHub_Trending/kh/khoj点击查看免费下载相关推荐Khoj 自托管怎么接入 Notion 工作区并索引笔记Khoj 自托管怎么接入 Notion 工作区并索引笔记 如果你的 Khoj 是自己部署的Docker 或 pip 安装在本机/服务器上想把 Notio人工智能AI 应用大模型RAG后端AI AgentKhoj 的 Markdown 知识库索引实战以 having_kids.markdown 为样例理解标题切分、嵌入与 Agent 检索Khoj 的 Markdown 知识库索引实战以 having_kids.markdown 为样例理解标题切分、嵌入与 Agent 检索 导读 本文以 kho人工智能AI 应用大模型RAG后端AI AgentKhoj Obsidian 插件接入指南在笔记库内搭建可对话、可检索的 AI 第二大脑Khoj Obsidian 插件接入指南在笔记库内搭建可对话、可检索的 AI 第二大脑 Khoj 是你的 AI 第二大脑它通过 Obsidian 社区插人工智能AI 应用大模型RAG后端AI Agent创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考