每日 Harness 开源 · Source
全部刊期 · All issues

每日 Harness

2026-09-21 · Monday, September 21, 2026

智能体工程化与可控性升级

视图 · View

今日重点 · Today's Highlights

agent-desktop2 - 用无障碍树和稳定引用替代像素猜测,渐进式遍历在密集桌面应用中可减少 78 - 96% token。

全文 ↓

Docling3 - 把 PDF 版面、表格、公式、OCR 及新加入的视频/XBRL 内容统一成可供生成式 AI 消费的结构化文档。

全文 ↓

mem04 - 以单次抽取、实体链接、混合检索和时间推理构建持久记忆,在 LoCoMo 上把分数从 71.4 推到 92.5。

全文 ↓

Openmsg5 - 让不同厂商 agent 把消息投递进彼此已有会话;跨人版本已通过 14 个规格验收案例,但互联网实战仍待验证。

全文 ↓

论文 · Papers

1 项 · 论文

本期重点An empirical study of harness design for coding agents1arxiv.org原文 ↗

arxiv.org

论文将 harness 视为可拆解的执行系统,在四个模型、SWE-Bench Verified 和 Terminal-Bench 2.1 上控制变量比较规划、动作空间与上下文策略。176 组设置显示,预算越紧,上下文管理越能避免 overflow;先做规则式省略再调用 LLM 摘要的效率最好,而可恢复省略没有带来准确率收益。规划对弱模型更像准确率支架,对强模型则主要节省成本;bash 能力强的模型用 bash-only 界面也能以更低成本完成命令行任务。

–

开源 / 项目 · Projects

12 项 · 开源 / 项目

AgentTrace6github.com原文 ↗

github.com

AgentTrace 把多步 agent 的工具调用、延迟和参数修复放进一套 SDK/仪表盘。Python 装饰器先以 Pydantic 校验 schema,失败时让 Groq 修复字符串/浮点混用、错误键名或缺失字段;FastAPI+SQLite 保存 trace,Next.js 提供 payload diff。它的价值在于把“偶发的工具参数幻觉”从致命异常变成可观察、可回放的运行时事件。

–

本期重点Openmsg5github.com原文 ↗

github.com

项目定义了跨 Claude Code、Codex、OpenCode 的会话内消息通道,消息不会另起进程,两个协作者也能跨机器通信。README 说明 0.1 已在同机 live session 测试,0.2 跨人实现通过 14 个 acceptance cases,但尚未有团队跑过互联网部署;这种状态把协议可行性与分布式运维风险清楚地区分开。

–

Emetgate7github.com原文 ↗

github.com

Emetgate 把“模型提议、内核验证”固化成源码写入门。真实会话会拒绝占位正文、转义错误、过期 hash 和越出 shadow copy 的路径,只让正确 body 进入提交流程;它针对编译成功但语义错误、模型忘记早先规则等问题,提供了比提示词更硬的边界。

–

meclaw8github.com原文 ↗

github.com

meclaw 用一个 Rust 二进制承载复杂 agentic system,参考实现 meclaw-os 把 voice cell、web cell 与 vault 放进按 cell 隔离的 kernel sandbox。系统同时兼容 OpenAI-compatible endpoint 和 MCP server,启动脚本可直接用 OpenRouter key 生成助手;“agent 组成 agentic OS”的抽象比单一编排库更宽,但也意味着安全模型必须覆盖整个运行时。

–

agent-term9github.com原文 ↗

github.com

AgentTerm 面向 Claude Code、Codex、Cursor CLI 等终端 agent,核心体验是区分、恢复并评论正在运行的会话,兼容 macOS 与 Windows/WSL。它把终端的文本输入输出保留下来,同时补上多会话管理和手机访问,而不是把 CLI 改造成封闭 IDE;因此更适合已有命令行习惯的团队试用。

–

System One Harness10github.com原文 ↗

github.com

该 harness 把有限动作编译成 typed questions,由 System One 返回动作概率,再按 read/write/destructive 风险阈值执行。内置订单示例五步完成、墙钟 0.98 秒、成本约 0.000208 美元;UHP core 40/40 conformance 说明协议层较完整,但作者明确只证明了小型确定性任务。

–

ZizkaDB11github.com原文 ↗

github.com

ZizkaDB 将 agent 决策写成 checksum-backed 审计链,提供 session replay、时间旅行调试、漂移检测、MCP 及 Python/TypeScript SDK。仓库以“从任一步回到根因”为产品入口,并把 Article 12 的留存要求映射到数据结构;这使它偏向合规和事后取证,而不只是普通日志库。

–

jevc12github.com原文 ↗

github.com

jevc 编译自然语言政策为确定性 verdict 程序:Jev 只回答窄证据问题并返回概率,最终结论由 reducer 在普通代码中计算。以禁止未经许可 commit 的规则为例,项目把会被模型反复遗忘的 Markdown 约束变成可测试门控;代价是策略必须能被拆成明确、可枚举的问题。

–

cleancode13github.com原文 ↗

github.com

CleanCode 用 canvas 记录每个变更的 worktree、agent、终端、服务、端口和依赖关系,节点可按就绪信号和退出码执行。下游任务会等待所有上游完成,服务失败则显式阻断后继;33 个 provider 和可复用模板把“重建开发上下文”的隐性劳动变成了可保存资产。

–

lgtm14github.com原文 ↗

github.com

lgtm 检查测试块与其实现是否真正对应,既能在 CLI/CI 运行,也能让 coding agent 通过 `/lgtm` skill 自动修复。项目在真实应用上处理过 15,000 多个测试块,并把其中 262 个逐一标注来拟合阈值;它关注的是绿色测试的有效性,而非再增加测试数量。

–

ColliePWA15colliepwa.dev原文 ↗

colliepwa.dev

ColliePWA 将 Herdr 会话聚合到手机端,保留完整 scrollback,按阻塞状态排序,并可在状态变化时发送 Web Push。网关默认只监听 localhost,再通过 Tailscale、WireGuard 或反向代理出网;多主机 crew 以 lead/deputy/peer 组织,每台机器保留自己的上传、journal、审计日志和故障转移状态。

–

Hush16github.com原文 ↗

github.com

hush 是带 abstention 的 GitHub issue/PR 分流 Action,只有标签概率和置信度同时过线才会写入标签。示例中 bug 100% 与 duplicate 97% 被接受,模糊案例 bug 72%/置信度 63% 则保持安静;单条处理耗时 202 - 530ms,说明“拒绝猜测”也被纳入了可测的产品行为。

–

行业动态 · Industry News

11 项 · 行业动态

Qwen Image 2.120qwen.ai原文 ↗

qwen.ai

Qwen 页面发布 Image 2.1 图像生成模型预览。该条的可核实信息是版本升级本身,读者应把它视作产品线更新,而非在缺少可比评测时推断能力跃迁。

–

Microsoft director: AI scraping “the largest theft of labor in human history”21tomshardware.com原文 ↗

tomshardware.com

纽约时报诉讼文件披露,微软内部曾把抓取称为“历史上最大的劳动盗窃”,并记录 Copilot 给 NYT 的点击率最多比 Bing 低 93%;OpenAI 内部则称 ChatGPT 对出版商构成生存威胁。材料把训练数据的 fair-use 抗辩和生成式产品对原站流量的替代效应放进同一法律事实链。

–

Step 5 Preview: Advancing the Pareto Frontier22stepfun.com原文 ↗

stepfun.com

StepFun 发布 Step 5 预览页,说明新模型的性能目标与预览方向。现阶段更适合把它看成路线节点,具体基准与可用范围仍应以正式发布材料为准。

–

How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip23spectrum.ieee.org原文 ↗

spectrum.ieee.org

IEEE 记录 Jalapeño 项目把内部 LLM 用在 RTL、验证和软件优化:DeepSeek 注意力内核从理论上限 0.31% 提升到 88.94%,约耗时 40 小时,矩阵乘单元面积又比优化的人类基线少 10%。后端布线和送厂主要由 Broadcom 承担,说明 AI 的收益集中在前端探索与首批芯片后的软件调优,而不是包办整个芯片流程。

–

KDE turns 30 and someone's brought an AI-native desktop proposal24theregister.com原文 ↗

theregister.com

KDE 30 周年 Akademy 的提案设想以加密、厂商无关的 Kadai 个人模型,为每位用户和设备动态编译 Plasma Activity。方案要求 declarative reconciliation、每组件 capability 与更丰富的 Activity metadata;它把 AI 放到桌面基础设施层,因此比侧边栏聊天更具架构野心,也更容易触发对主权和可控性的争论。

–

Tin: full-text search for Postgres25planetscale.com原文 ↗

planetscale.com

PlanetScale 的 Tin 用 Postgres ctid 作为 posting 标识,借现代 CPU 做向量化交并;其词组查询达到 242 QPS、p99 212ms,Wikipedia COUNT(*) 更达到 10,260 QPS、2ms p99。作者称多数场景吞吐至少高于替代方案 8 倍,关键不在再造一个独立搜索集群,而是把索引和现有 Postgres 存储路径贴近。

–

How Notion handles concurrent editing with CRDTs27notion.com原文 ↗

notion.com

Notion 在 CRDT 中把 text instance 与 block 分离映射,并用 L/R search label 追踪拆分后的文本片段;原本可能要加载 100 个 block,标签能把查询缩到一个,紧凑编码还带来约五倍存储效率。这个设计承认并发编辑的困难不只在合并顺序,也在服务端如何低成本定位被切碎的文本。

–

博客文章 · Blog Posts

13 项 · 博客文章

Quoting voxium28simonwillison.net原文 ↗

simonwillison.net

Simon Willison 转述 voxium 开发者对公司内部大量采用 Claude Code 的观察,焦点是 agent 从个人试验走向组织日常后的工作方式变化。条目更像一份现场经验切片,价值在于记录采用过程中的组织信号,而非宣布一项新模型能力。

–

Deterministic Core, Non-Deterministic Shell29outdata.net原文 ↗

outdata.net

文章主张把状态转换、业务规则和可复现检查放在确定性核心,让模型负责探索、解释与调度等非确定性外壳。这样的分层能把“模型偶尔答错”限制在可替换区域,同时保留传统测试对核心状态机的约束。

–

Adversarial examples for fast hash functions30thomasahle.com原文 ↗

thomasahle.com

作者不是用随机碰撞展示哈希风险,而是专门构造针对快速哈希的对抗输入。这个角度把性能优化后的分布假设暴露出来:在攻击者可选输入的场景,平均速度和最坏情况需要分开评估。

–

Software Sandboxing: The Basics (2025)31blog.emilua.org原文 ↗

blog.emilua.org

Emilua 将沙箱归纳为程序化、无管理员权限的可裁量权限下降,建议用进程隔离、actor 消息和 capability 文件描述符组合出边界。作者偏爱 Capsicum 的 cap_enter,因为它关闭 ambient authority;引用的 Chromium 对比中,Capsicum 约 100 行改造即可工作,而 seccomp 方案约 11,301 行,凸显接口复杂度差异。

–

Prompts aren’t Real32evaluation.club原文 ↗

evaluation.club

文章把 prompt 当作不稳定的软件接口,指出自然语言文本本身不是可靠的行为规格。它因此把评测重点移向可重复的输入、上下文和输出契约:只有能被系统化测试的行为,才配称作接口。

–

Trying the Software Factory Pattern33lethain.com原文 ↗

lethain.com

这次实验把需求、实现、测试和交付拆成可重复的流水线单元,观察角色边界和反馈节奏如何改变吞吐与返工。其启发不是机械套用工厂比喻,而是把软件过程中的隐性等待与交接显式化。

–

AI and the Destruction of the Creative Commons34chesterwisniewski.com原文 ↗

chesterwisniewski.com

文章认为生成式 AI 在吸收 Creative Commons 内容时,可能绕开署名、访问和回馈机制,削弱原本支撑开放内容的激励循环。论点把版权合法性之外的生态基础设施摆到前景:内容可训练,不代表贡献者仍能持续生产。

–

Chat-based Large Language Models replicate the mechanisms of a psychic's con35softwarecrisis.dev原文 ↗

softwarecrisis.dev

“LLMentalist”按冷读流程拆解聊天模型:先用模糊开场建立连接,再通过用户反应缩小目标,最后依靠主观验证让泛化陈述显得精准。它解释了流畅、迎合和用户自我投射如何共同制造“被理解”的体验,但这不是事实获取能力的证据。

–

Why MCP Was Always a Bad Idea37maharship.com原文 ↗

maharship.com

这篇批评从协议边界、工具发现和状态语义入手,认为 MCP 把过多复杂性转嫁给客户端与模型。它适合作为架构反例阅读:即便工具接入变得标准化,权限、生命周期和错误语义仍必须由系统设计者明确承担。

–

Orchestrating Claude Code Agents: The Chief of Staff Pattern38asyncdot.com原文 ↗

asyncdot.com

chief-of-staff 模式用一个总协调会话拆解目标、分派任务、汇总结果,再让多个 Claude Code 会话各自执行。关键不在并发数量,而在委派、回报和全局上下文的显式协议;否则多个 agent 只会把未整理的状态互相放大。

–

Telling a Computer to Do Things39will-keleher.com原文 ↗

will-keleher.com

作者从单命令走到带条件、并发和循环的 shell,再把这条能力链连接到 agent。npm 失败处理和“测试直到失败”的例子说明,真正的自动化依赖对程序组合的理解,而不是 GUI 是否提供一个现成按钮。

–

What Zig felt like, coming from Rust40besok.github.io原文 ↗

besok.github.io

Rust 转 Zig 的 JSONPath 实作暴露了两种语言的抽象差异:Rust 倾向 immutable/monadic transformation,Zig 则让作者直接操作 allocator、cursor 和列表。两者都能表达递归与 sum type,但内存管理责任的转移会重塑 API 和算法写法。

–

引用来源 · References

48 条 · 引用
  1. 1 An Empirical Study of Harness Design for Coding Agents. arXiv:2609.20804https://arxiv.org/abs/2609.20804 ↩ 回到正文 · back to text
  2. 2 agent-desktophttps://github.com/lahfir/agent-desktop ↩ 回到正文 · back to text
  3. 3 Doclinghttps://github.com/docling-project/docling ↩ 回到正文 · back to text
  4. 4 mem0https://github.com/mem0ai/mem0 ↩ 回到正文 · back to text
  5. 5 Openmsghttps://github.com/marciob/openmsg ↩ 回到正文 · back to text
  6. 6 AgentTracehttps://github.com/mohitkumar188/AgentTrace ↩ 回到正文 · back to text
  7. 7 Emetgatehttps://github.com/emetgate/emetgate ↩ 回到正文 · back to text
  8. 8 meclawhttps://github.com/mmeyerlein/meclaw/ ↩ 回到正文 · back to text
  9. 9 agent-termhttps://github.com/albertwujj/agent-term ↩ 回到正文 · back to text
  10. 10 System One Harnesshttps://github.com/HarnessRouter/SystemOneHarness ↩ 回到正文 · back to text
  11. 11 ZizkaDBhttps://github.com/ZIZKA-AI-SL/ZizkaDB ↩ 回到正文 · back to text
  12. 12 jevchttps://github.com/doronp/jevc ↩ 回到正文 · back to text
  13. 13 cleancodehttps://github.com/chen-985211/cleancode ↩ 回到正文 · back to text
  14. 14 lgtmhttps://github.com/stardeckai/lgtm ↩ 回到正文 · back to text
  15. 15 ColliePWAhttps://colliepwa.dev/ ↩ 回到正文 · back to text
  16. 16 Hushhttps://github.com/emreozyoruk/hush ↩ 回到正文 · back to text
  17. 17 OpenAI's Sam Altman to Brief UN Security Council Next Weekhttps://www.reuters.com/business/openais-sam-altman-to-brief-un-security-council-next-week-during-2026-09-18/ ↩ 回到正文 · back to text
  18. 18 Samsung is expected to more than double output of its HBM4 and HBM4E DRAMhttps://en.sedaily.com/finance/2026/09/20/samsung-to-double-hbm4-output-next-year-sources-say ↩ 回到正文 · back to text
  19. 19 ChatGPT now knows what you do on other websites via ad collectorhttps://www.buchodi.com/chatgpt-now-knows-what-you-do-on-other-websites-via-ad-collector/ ↩ 回到正文 · back to text
  20. 20 Qwen Image 2.1https://qwen.ai/blog?id=qwen-image-2.1 ↩ 回到正文 · back to text
  21. 21 Microsoft director: AI scraping “the largest theft of labor in human history”https://www.tomshardware.com/tech-industry/artificial-intelligence/microsoft-director-called-ai-scraping-the-largest-theft-of-labor-in-human-history-while-openai-head-brands-chatgpt-an-existential-threat-to-publishers-revelations-come-from-legal-briefs-filed-in-nyt-lawsuit ↩ 回到正文 · back to text
  22. 22 Step 5 Preview: Advancing the Pareto Frontierhttps://www.stepfun.com/step-5-preview ↩ 回到正文 · back to text
  23. 23 How OpenAI Used Its Own LLMs to Design Its Jalapeño Chiphttps://spectrum.ieee.org/llms-for-chip-design ↩ 回到正文 · back to text
  24. 24 KDE turns 30 and someone's brought an AI-native desktop proposalhttps://www.theregister.com/software/2026/09/18/kde-turns-30-and-someones-brought-an-ai-native-desktop-proposal/5297282 ↩ 回到正文 · back to text
  25. 25 Tin: full-text search for Postgreshttps://planetscale.com/blog/introducing-tin ↩ 回到正文 · back to text
  26. 26 The Hugging Face Hack Wasn't What It Was Cracked Up to Behttps://www.wsj.com/opinion/the-hugging-face-hack-wasnt-what-it-was-cracked-up-to-be-e00cf3fa ↩ 回到正文 · back to text
  27. 27 How Notion handles concurrent editing with CRDTshttps://www.notion.com/blog/how-notion-handles-concurrent-editing-with-crdts ↩ 回到正文 · back to text
  28. 28 Quoting voxiumhttps://simonwillison.net/2026/Sep/20/voxium/ ↩ 回到正文 · back to text
  29. 29 Deterministic Core, Non-Deterministic Shellhttps://outdata.net/blog/260803 ↩ 回到正文 · back to text
  30. 30 Adversarial examples for fast hash functionshttps://thomasahle.com/blog/adversarial-examples-for-hashes/ ↩ 回到正文 · back to text
  31. 31 Software Sandboxing: The Basics (2025)https://blog.emilua.org/2025/01/12/software-sandboxing-basics/ ↩ 回到正文 · back to text
  32. 32 Prompts aren’t Realhttps://evaluation.club ↩ 回到正文 · back to text
  33. 33 Trying the Software Factory Patternhttps://lethain.com/software-factory-experiment/ ↩ 回到正文 · back to text
  34. 34 AI and the Destruction of the Creative Commonshttps://www.chesterwisniewski.com/post/2026-09-13-ai-is-destroying-the-creative-commons/ ↩ 回到正文 · back to text
  35. 35 Chat-based Large Language Models replicate the mechanisms of a psychic's conhttps://softwarecrisis.dev/letters/llmentalist/ ↩ 回到正文 · back to text
  36. 36 Self-hosted inference orchestrators compared: LocalAI, exo, GPUStack, vLLMhttps://www.nexlab.net/articles/self-hosted-inference-orchestrators-compared-2026/ ↩ 回到正文 · back to text
  37. 37 Why MCP Was Always a Bad Ideahttps://maharship.com/blog/why-mcp-was-always-a-bad-idea/ ↩ 回到正文 · back to text
  38. 38 Orchestrating Claude Code Agents: The Chief of Staff Patternhttps://asyncdot.com/blog/chief-of-staff-pattern-orchestrating-claude-code-sessions/ ↩ 回到正文 · back to text
  39. 39 Telling a Computer to Do Thingshttps://will-keleher.com/posts/telling-your-computer-to-do-things/ ↩ 回到正文 · back to text
  40. 40 What Zig felt like, coming from Rusthttps://besok.github.io/posts/what-zig-felt-like-coming-from-rust/ ↩ 回到正文 · back to text
  41. 41 supabase/supabasehttps://github.com/supabase/supabase ↩ 回到正文 · back to text
  42. 42 huggingface/transformershttps://github.com/huggingface/transformers ↩ 回到正文 · back to text
  43. 43 higgsfield-ai/higgsfieldhttps://github.com/higgsfield-ai/higgsfield ↩ 回到正文 · back to text
  44. 44 cloudflare/quichehttps://github.com/cloudflare/quiche ↩ 回到正文 · back to text
  45. 45 ruanyf/weeklyhttps://github.com/ruanyf/weekly ↩ 回到正文 · back to text
  46. 46 yynxxxxx/Codex-Xhttps://github.com/yynxxxxx/Codex-X ↩ 回到正文 · back to text
  47. 47 rui314/moldhttps://github.com/rui314/mold ↩ 回到正文 · back to text
  48. 48 NVIDIA/TensorRT-LLMhttps://github.com/NVIDIA/TensorRT-LLM ↩ 回到正文 · back to text