All collections
30 connected reading paths. Start with any collection, then follow it in order.
LVM 系列
2 articlesView all articles
LobeChat 部署系列
2 articlesView all articles
iSCSI专题系列
1 articlesPython数据分析系列
4 articlesView all articles
vLLM 高性能推理系列
4 articlesView all articles
Docker 容器化系列
3 articlesView all articles
Python 数据科学系列
5 articlesView all articles
Ollama 部署系列
2 articles机器学习基础系列
45 articlesView all articles
- 学习资源
- 核心概念
- 线性代数基础
- 概率统计基础
- 微积分基础
- 损失函数与优化
- 线性回归
- 多项式回归与正则化
- 逻辑回归
- 决策树
- 随机森林
- 支持向量机
- K近邻算法
- 朴素贝叶斯
- 聚类算法
- 降维技术
- 模型评估指标
- 交叉验证
- 过拟合与欠拟合
- 超参数调优
- 神经网络基础
- 前馈神经网络
- 激活函数
- 反向传播
- 优化算法详解
- 批归一化
- Dropout详解
- 卷积神经网络
- 循环神经网络
- 注意力机制
- Transformer
- 迁移学习
- 实战项目
- 集成学习
- XGBoost与LightGBM
- 自编码器
- 生成对抗网络
- 图神经网络
- 混合专家模型
- 强化学习基础
- 自监督学习
- 知识蒸馏
- LoRA与参数高效微调
- 模型量化与剪枝
- 机器学习基础系列——梯度提升
LLM 应用开发系列
39 articlesView all articles
- Embedding 嵌入向量
- 向量数据库
- RAG 检索增强生成
- Prompt Engineering
- LangChain 框架
- Agent 智能体
- Function Calling
- 部署与优化
- 安全防护
- 对话记忆管理
- Streaming 流式输出
- LlamaIndex 框架
- 结构化输出
- 多模态 LLM 应用
- LLM 评估与测试
- 多模型路由
- 缓存策略
- 成本优化
- 监控与可观测性
- 高级 RAG 技术
- GraphRAG 知识图谱
- 长上下文处理
- LLM 微调实战
- 本地模型部署
- MCP 协议开发
- AI 搜索引擎
- 代码生成与执行
- 文档智能处理
- 语音交互系统
- 多租户架构
- Prompt 版本管理
- 混合检索策略
- 对话系统设计
- 内容审核与过滤
- 个性化推荐
- 知识库构建
- AI 写作助手
- 实时协作
- 边缘部署
企业级RAG应用系列
4 articles前端开发者的 AI 进阶之路
15 articlesView all articles
- 01. AI 入门:前端开发者的第一课
- 02. 调用大模型 API:从 fetch 到流式响应
- 03. Prompt 工程:让 AI 听懂你的话
- 04. Embedding 与向量:把文字变成数字
- 05. RAG:让 AI 回答你的专属数据
- 06. Function Calling:让 AI 调用你的 API
- 07. Agent 入门:从对话到自主行动
- 08. MCP 协议:Agent 的通用接口
- 09. 构建你的第一个 AI 聊天应用
- 10. 进阶方向:多模态、本地部署与成本优化
- 番外 1. LLM 简史:从 Transformer 到今天
- 番外 2. API 参数全解:temperature 到 logprobs
- 番外 3. Token 的秘密:BPE、中文更贵、怎么省
- 番外 4. 上下文窗口进化史:KV Cache 与长上下文
- 番外 5. 幻觉与开源崛起:LLM 的两个关键议题
Python 新手的 AI 进阶之路
28 articlesView all articles
- 01. AI 入门:Python 新手的第一课
- 02. 调用大模型 API:从 requests 到流式响应
- 03. Prompt 工程:让模型稳定输出你要的结果
- 04. 结构化输出:用 Pydantic 把 LLM 变成稳定的函数
- 05. Embedding 与向量:把文字变成数字
- 06. RAG 实战:搭建一个能回答你文档的本地知识库
- 07. Function Calling:让 LLM 调用你的 Python 函数
- 08. Agent 入门:从工具调用到自主循环
- 09. MCP 协议:Agent 的通用接口
- 10. 本地模型:用 Ollama 跑开源 LLM
- 11. 用 Streamlit 构建 AI 聊天应用
- 12. 进阶方向:评测、监控、成本与微调
- 番外 1. LLM 简史:从 Transformer 到 2026
- 番外 2. API 参数全解:temperature 到 logprobs
- 番外 3. Token 的秘密:BPE、中文为什么更贵
- 番外 4. 上下文窗口进化史与 KV Cache
- 番外 5. 幻觉与开源崛起
- 番外 6. 异步与并发:批量调用 LLM 的工程模式
- 番外 7. 提示词注入与防御
- 番外 8. 多模态入门:图像、语音与视频
- 番外 9. Embedding 的来龙去脉:从 one-hot 到向量空间
- 番外 10. Embedding 演化实战:从 TF-IDF 到 LLM 的 token 表
- 番外 11. 从 One-Hot 到 LLM:NLP 三十年的完整演化
- 番外 12. Embedding 模型选型实战
- 番外 13. 文本切分策略与 chunk_size 选择
- 番外 14. 向量数据库 Qdrant、混合检索与 Reranker
- 番外 15. RAG 评估方法:把感觉变成数字
- 番外 16. uvicorn --reload 与本地大模型的相处难题
Agent 模式手册:从最小循环到生产可用
23 articlesView all articles
- 01. Agent 是什么:从一次性 Function Call 到自主循环
- 02. ReAct:把 Agent 的思考显式化
- 03. 工具设计的艺术:粒度、命名、错误反馈与工具爆炸
- 04. Plan-and-Execute:先规划,再执行,失败时重规划
- 05. Reflection:让 Agent 自己挑毛病、自己改
- 06. Tree of Thoughts:从线性循环到搜索式推理
- 07. 记忆系统:短期对话、长期事实、过程记忆
- 08. 上下文工程:Agent 最难的不是推理,是喂什么进去
- 09. 工具调用协议的分裂:OpenAI、Anthropic、Hermes
- 10. Hermes 3 实战:在 Ollama 上跑一个本地工具调用 Agent
- 11. 开源 Agent 全家桶:Hermes 3 + 向量库的离线 RAG-Agent
- 12. 多 Agent 协作:Orchestrator、Debate、Handoff 与它们的陷阱
- 13. Code Execution Agent:让 Agent 真的能写代码、跑代码
- 14. Browser & Computer Use:让 Agent 看屏幕、点鼠标
- 15. Human-in-the-Loop:什么时候让 Agent 停下来问人
- 16. 可观测性与评测:Agent 的 trace、replay 与回归测试
- 17. 生产化:上下文压缩、成本、延迟、失败恢复、安全边界
- 18. 框架世界观对比:Claude Agent SDK、LangGraph、CrewAI、Pydantic AI
- 19. LangGraph 状态机编排:让 Agent 的流程真正可控
- 20. LlamaIndex 数据侧深度:索引、检索、合成的分工
- 21. Benchmark 全景:SWE-bench、AgentBench、τ-bench 怎么读
- 22. Agent 安全边界:Prompt Injection、权限最小化、沙箱逃逸
- 23. Agent 失败模式分类学:为什么它总在你不期待的地方崩
Code as Agent Harness:代码是 Agent 的脊梁
20 articlesView all articles
- 01. 代码是推理的外包地:Program of Thoughts 与 Chain of Code
- 02. 代码还是文字?CodeSteer 与推理模式的动态选择
- 03. 从执行反馈中学习:CodeRL 与迭代自修正
- 04. Voyager:用代码探索开放世界的终身学习 Agent
- 05. CodeAct:把代码执行变成通用 Agent 的动作空间
- 06. Self-Debugging:让模型自己读报错、自己改代码
- 07. SWE-agent:把真实 GitHub Issue 变成 Agent 的考场
- 08. Reflexion:用语言反思替代梯度更新
- 09. LATS:把树搜索和语言反思结合起来
- 10. AlphaCodium:用测试驱动的迭代流程生成代码
- 11. AgentCoder:拆分角色,让代码生成和测试设计分开
- 12. MapCoder:用类比和检索辅助代码生成
- 13. InterCode:把交互式编程环境变成强化学习的训练场
- 14. 代码 Agent 的设计模式:前 13 篇的横向梳理
- 15. 代码 Agent 的安全边界:沙箱、权限与提示注入
- 16. 大代码库的上下文管理:Agent 如何在百万行代码里找路
- 17. 代码 Agent 评测方法的设计反思
- 18. 多模态代码 Agent:视觉输入加入代码生成的工作流
- 19. Agent 框架横向对比:LangGraph、AutoGen、CrewAI、OpenHands
- 20. 模型自身能力对 Agent 表现的影响
Claude Code 实战手册(从 WSL 到进阶)
10 articlesView all articles
AI 寓言集:用故事讲透五个核心概念
20 articlesView all articles
- 01. 图书馆里的低语(Self-Attention)
- 02. 雪夜下山的瞎子(Gradient Descent)
- 03. 完美的临摹学徒(Overfitting 与正则化)
- 04. 把世界搬上书架的图书管理员(Embedding)
- 05. 学了三年突然顿悟的少年(Grokking 与涌现)
- 06. 在雾中看潮水的渔夫(Bayes 推断)
- 07. 十口井的旅人(Multi-armed Bandit)
- 08. 同一双眼睛走遍全村(Convolution / CNN)
- 09. 学新手艺就忘旧的画师(Catastrophic Forgetting)
- 10. 画赝品的人和他的宿敌(GAN)
- 11. 一张照片骗过所有人的眼睛(Adversarial Examples)
- 12. 翻山送信的邮差(Residual Connection)
- 13. 把老匠人装进一个孩子的梦里(知识蒸馏)
- 14. 山神的三个承诺(Reward Hacking / Goodhart)
- 15. 守门人和十个专家(Mixture of Experts)
- 16. 天生注定的那张彩票(Lottery Ticket Hypothesis)
- 17. 能把尘土还原成瓷瓶的人(Diffusion)
- 18. 够大的雨终会灌满湖(Scaling Laws / Emergence)
- 19. 会大声自言自语的棋手(Chain of Thought)
- 20. 一支笔里藏着五种颜色(Superposition / 机械可解释性)
实图小记:不写一行代码做一个 iOS app(连载中)
3 articlesAgent 评测手册:怎么判断一个 Agent 好不好
1 articlesLLM 进阶系列
7 articles模型是怎么训练出来的:从零写一个迷你 GPT
8 articles向量数据库系列
8 articles模型推理优化系列
8 articlesLLM 工程路线图
41 articlesView all articles
- LLM 工程路线图:从概念理解到可运行系统
- 项目 1-2:Tokenizer、词表与 Embedding
- 项目 3-8:位置、注意力与 Transformer Block
- 项目 9-16:Decoding、KV Cache、长上下文与推理系统
- 项目 17-33:MoE、数据、后训练、评测、RAG、Agent 与安全
- 项目 34:十二周执行计划与 Capstone
- 34 个 LLM 工程项目验收清单
- 从零实现 tokenizer
- embedding 与语义几何
- 位置编码(learned / sinusoidal / RoPE / ALiBi)
- 手写 scaled dot-product attention
- 扩展到多头注意力机制(Multi-Head Attention)
- 构建完整的 Transformer Decoder Block
- Mini-former 训练实战:从随机扰动到文本预测
- 训练目标对比:Causal、Masked 与 Prefix LM
- 概率的剪裁:解码策略与采样 Dashboard
- 打破串行咒语:投机解码(Speculative Decoding)
- 显存的吞噬者:KV Cache 机制与显存预算
- 架构的进化:MQA、GQA 与 MLA 深度解析
- 跨越万词鸿沟:长上下文(Long Context)的系统性挑战与解法
- 频率的炼金术:RoPE Scaling 与长度外推(Extrapolation)
- IO 感知的艺术:FlashAttention 的硬件级优化
- 硬件精算:显存带宽、算力与硬件预算(Hardware Budget)
- 稀疏性的调度艺术:实现双专家 MoE 路由(MoE Router)
- 计算的权衡:稠密(Dense)与稀疏(MoE)模型的全方位对比
- 线性序列的回归:状态空间模型(SSM)与线性注意力
- 文本的扩散:扩散语言模型(Diffusion LM)初探
- 文明的数字工业化:构建大规模预训练数据管线
- 自我进化的循环:合成数据(Synthetic Data)的生成、过滤与证明
- 预测未来的算力账本:缩放法则(Scaling Laws)与曲线拟合
- 灵魂的对齐:SFT、指令微调与偏好优化(DPO)
- 理性的驯化:从 RLHF、PPO 到 GRPO 的进化史
- 数字的炼金术:模型量化(Quantization)深度解密
- 吞吐量之王:Serving Stack 与推理引擎横评
- 拒绝虚假繁荣:构建严谨的模型评测(Evaluation Harness)
- 检索增强生成的艺术:工业级 RAG 架构拆解
- 构建 Tool Use 与 Agent Loop
- 多模态视觉语言桥接器(Vision-Language Adapter)
- 打开黑盒:机械可解释性(Mechanistic Interpretability)
- 防线与红队:构建 LLM 安全评估体系
- 终极试炼:构建你的私有化大模型系统(Capstone Project)
深入理解 AI Agent · Harness 工程笔记
11 articlesAgent 运行时实验室:从抓包到可重放系统
18 articlesView all articles
- 00. 导读:从一次抓包进入 Agent 运行时
- 01. 系统提示词是怎样被组装出来的
- 02. Prompt 长度、上下文增长与缓存成本
- 03. Prompt 注入发生在哪一层
- 04. 工具 Schema 怎样改变 Agent 的行为
- 05. 模型调用工具之后,运行时发生了什么
- 06. 工具结果应该怎样回填给模型
- 07. 并行工具调用真的更快吗
- 08. Agent 为什么需要状态机
- 09. Agent 权限不是 allow 和 deny 两个值
- 10. 如何防止有副作用的操作重复执行
- 11. Agent 的错误应该在哪一层处理
- 12. 上下文、摘要和长期记忆不是一回事
- 13. 多 Agent 的价值不是同时运行多个模型
- 14. 怎样记录一条可以重放的 Agent 轨迹
- 15. 怎样评测一个 Agent,而不是评价最终回答
- 16. Agent 上线前怎样设计回归测试集
- 17. 怎样给 Agent 做一次故障演练
Agent Runtime Lab: From Packet Capture to Replay
18 articlesView all articles
- 00. Reading the Agent Runtime from Its Traces
- 01. How System Prompts Are Assembled
- 02. Prompt Length, Context Growth, and Cache Cost
- 03. Where Prompt Injection Enters an Agent
- 04. How Tool Schemas Change Agent Behavior
- 05. What Happens After the Model Calls a Tool
- 06. How Tool Results Should Return to the Model
- 07. Are Parallel Tool Calls Actually Faster?
- 08. Why an Agent Needs a State Machine
- 09. Agent Permission Is More Than Allow or Deny
- 10. Preventing Duplicate Side Effects
- 11. Which Layer Should Handle an Agent Error?
- 12. Context, Summaries, and Long-Term Memory Are Different
- 13. Multi-Agent Value Is Not Multiple Models Running
- 14. Recording a Replayable Agent Trace
- 15. Evaluating an Agent Beyond Its Final Answer
- 16. Building an Agent Regression Suite Before Release
- 17. Running a Fault-Injection Exercise for an Agent
从灾备走向数据中心建设
1 articlesFrom Disaster Recovery to Data Center Infrastructure
1 articlesRecent writing
Back to collections ↑17. Running a Fault-Injection Exercise for an Agent
A successful happy path proves that the system can run. A controlled failure shows whether it knows where to stop and how to recover.
16. Building an Agent Regression Suite Before Release
A regression suite is not a rerun of a few demos. It is an executable record of failures the system must not repeat.
09. Agent Permission Is More Than Allow or Deny
Authorization decides whether one principal may perform one action on one resource under specific conditions
07. Are Parallel Tool Calls Actually Faster?
Concurrency shortens only independent waiting time and adds result joining, cancellation, and resource pressure
12. Context, Summaries, and Long-Term Memory Are Different
Longer retention does not create stronger knowledge; unverified memory carries old errors into new tasks
15. Evaluating an Agent Beyond Its Final Answer
A plausible answer does not prove that evidence is real, execution was authorized, or interruption recovery is correct
01. How System Prompts Are Assembled
The model does not receive one prompt. It receives context blocks assembled from different sources for different purposes.
06. How Tool Results Should Return to the Model
The model does not observe local execution; the runtime must carry evidence into the next turn
04. How Tool Schemas Change Agent Behavior
The model does not operate the system directly; it selects from the actions exposed by the runtime
10. Preventing Duplicate Side Effects
Retrying a model call spends computation; retrying a send, payment, or create operation can change the external world twice