每日 Harness 开源 · Source
全部刊期 · All issues

每日 Harness

2026-09-07 · Monday, September 7, 2026

智能代理工程化与治理升温

视图 · View

今日重点 · Today's Highlights

Harnessing the Universal Geometry of Embeddings1 - 论文提出无需配对样本即可把文本嵌入从一个向量空间翻译到另一个空间的方法,并证明不同架构间存在可利用的普适语义几何;同一结构也让只拿到向量的攻击者能够推断敏感属性。

全文 ↓

indic-modernBERT3 - 这是从头训练的 Hindi-first 编码器,188M 参数、8192 上下文、约 28B Hindi tokens,单张 RTX 4090 用时五天;它给低资源语言提供了可复现实验基线。

全文 ↓

Wayfinder4 - 项目把非确定性应用的评估拆成数据集、rubric、trace、回归门和在线实验,示例是逐步演化的航班搜索应用;重点在让“可靠”变成可重复运行的工程流程。

全文 ↓

OKF Agent Memory5 - Git-native 的 Markdown/YAML 记忆库用 BM25 和图校验支撑 coding agent 的搜索前写入,实测检索低于 300µs、图校验约 4ms,并宣称可减少约 80% token 消耗。

全文 ↓

论文 · Papers

2 项 · 论文

本期重点Harnessing the Universal Geometry of Embeddings1arxiv.org原文 ↗

arxiv.org

作者构造了一个只依赖未配对向量的空间对齐流程,不需要访问原始编码器或训练匹配数据;跨模型的高余弦相似度表明语义结构比具体坐标更稳定。反面结果同样重要:向量服务暴露的表示可以被用来做分类和属性推断,因此嵌入发布应被当作潜在隐私接口。

–

本期重点Large-Language Models as a Cognitive Virus2arxiv.org原文 ↗

arxiv.org

论文将 LLM 的认知扩散写成带传播、康复和强化项的状态转换模型,而非简单的“用户是否使用”二元统计。模型预测在临界区域会出现 tipping point,采用率的微小变化可能把人群推入持久耦合;作者提出降低传播率、提高可逆性作为“认知免疫”方向,但尚未给出真实社会实验验证。

–

开源 / 项目 · Projects

15 项 · 开源 / 项目

Mador6github.com原文 ↗

github.com

Mador 用约 80 行运行时把 Proxy 状态元组接到任意 DOM:`r(selector, update, read)` 记录依赖,状态变化后只更新相关节点。没有组件、模板或虚拟 DOM,因而适合把微型交互直接嵌进现有页面;代价是开发者要自行管理结构和生命周期。

–

MaskShift7github.com原文 ↗

github.com

这是一个本地优先的终端 agent harness,把工具 schema 写进提示,再解析模型文本中的调用;Node 22 内置能力即可运行,内置 149 个工具、44 个 skills,支持 Ollama、OpenAI、Anthropic、Gemini 和 vLLM。检查点、审计轨迹与延迟加载 MCP 让长任务更容易回放,但“最大化”工具面也扩大了权限治理范围。

–

Byosynch8byosynch.com原文 ↗

byosynch.com

Byosynch 以 SSH 直连用户自己的存储,在 macOS 与 Linux 间持续双向同步,事件监听和定期扫描共同覆盖遗漏变更,并保留冲突文件。私钥和文件不经过中转服务器是清晰的信任边界;它不提供端到端加密,个人版试用 30 天后为每年 49 美元,这两点决定了适用场景。

–

Agentic OS9github.com原文 ↗

github.com

meclaw 将目录树当作运行时拓扑:每个文件夹是 actor,路径是消息路由,Rust 单二进制和 SQLite 保存状态,实体可拥有独立沙箱。README 记录首次冷启动约 40 秒、25MB 二进制,之后低于 1 秒;把拓扑放进文件/Git 便于审计,但也把文件变更提升成系统级行为。

–

toolog10github.com原文 ↗

github.com

toolog 将 Claude Code transcript tailer 与 OTLP receiver 通过 `tool_use_id` 汇入 SQLite WAL/FTS5,再用确定性规则和本地模型给调用打风险标签。监听仅绑定 loopback,CI 模式还能盘点 socket;这种“先完整留痕、后做判断”的结构比只收集最终 diff 更能定位代理的越权链条。

–

ZeroThesis11zerothesis.com原文 ↗

zerothesis.com

ZeroThesis 的设想是让多个 agent 围绕公开问题共享实验历史、提交方案,并在隔离沙箱中复现实验。它把研究协作的记忆和复现步骤放在同一平台,核心挑战转为实验 provenance、资源配额和结果可比性;项目页面目前主要展示这一工作流愿景。

–

HakoVM12github.com原文 ↗

github.com

HakoVM 利用 macOS Apple Containerization 构造 Firecracker 风格微型虚拟机,OCI 启动约 300ms、exec p50 约 8ms、常驻约 20MB RSS,且默认没有网卡、守护进程或 root。它支持提交/克隆工作区、MCP 和 Claude hook,把一次 agent 操作封装成可丢弃的内核边界。

–

本期重点indic-modernBERT3github.com原文 ↗

github.com

README 给出 22 层、8192 context 和新的 Hindi BPE;在 NER 上 F1 为 0.8001,MASSIVE 为 0.4731,检索任务在 DPR 之后最强。数据配方把 Sangraha 23.6B 与 IndicCorp V2 4.85B token 合并,展示了专门化语料对编码器能力的直接影响。

–

本期重点Wayfinder4github.com原文 ↗

github.com

Wayfinder 的参考实现用规则 judge、人类 judge、LLM judge 和在线评估逐层比较同一航班搜索系统,所有运行都保存 trace 和回归结果。它把提示或模型升级视为实验,而不是一次性发布;因此能具体观察可靠性、成本与用户成功率的权衡。

–

Twinrun13github.com原文 ↗

github.com

Twinrun 对 base/head 代码使用相同输入并比较行为差异,无需预先写规格;默认 24 个探针、20 秒超时、每个样本重复两次。项目在 509 个 click 提交、1630 个 callable 上发现 243 个差异,其中捕获了后来回滚的 7 个变更中的 5 个,说明差分测试适合补足规格空白。

–

本期重点OKF Agent Memory5github.com原文 ↗

github.com

OKF v0.2 是纯 Go 单二进制,知识以 Git 可审查的 Markdown/YAML 存放,提供 MCP、渐进披露和 search-before-write 约束。README 报告内存占用低于 15MB、图校验约 4ms;这让记忆层更像可维护的知识库,而不是不可见的向量黑箱。

–

dsnitch14github.com原文 ↗

github.com

dsnitch 在 cgroup v2 上用 eBPF 归因 Docker 容器的 TCP、UDP、ICMP 与 DNS 出口,Hickory 负责 CNAME 解码,TUI 可实时查看,`-s` 则输出脚本格式。项目声称 CPU 开销低于 1%,要求 Linux 5.8+ 和 BTF;它提供的是连接级证据,不能替代应用层请求审计。

–

vernLLM15github.com原文 ↗

github.com

vernLLM 把多提供商调用封装在进程内,提供重试、熔断、fallback、限流、缓存和 middleware。示例配置为 500 RPM、100k TPM、20 并发、超时 10 秒和最多 3 次重试;这些参数把供应商波动显式化,但缓存一致性和跨提供商语义差异仍需由业务定义。

–

webmcp-stack16github.com原文 ↗

github.com

webmcp-stack 从 OpenAPI 或验证 schema 生成 WebMCP 工具,并标注只读、写入和破坏性操作;写操作默认 withheld,必须经过确认,还能标记 untrusted 与 PII 字段。示例生成超过 70 个工具,CI 会验证 schema,核心贡献是把 agent 的权限审查前移到代码生成阶段。

–

XHFS17codeberg.org原文 ↗

codeberg.org

XHFS 提供带并发读写保证、内置 WebDAV 的 Copy-on-Write 文件系统。它把一致性读写和 WebDAV 组合在同一方案中。

–

行业动态 · Industry News

14 项 · 行业动态

Following legal advice, the Nitter project will continue18github.com原文 ↗

github.com

Nitter 团队在 2026 年 8 月 24 日收到法律函件后,依据法律咨询决定继续维护。项目仍以 Nim + Redis/Valkey 代理公开时间线,不加载 JavaScript 或广告;README 给出的典型页面约 60KB、速度可达原站的 2 - 4 倍,后续重点会落在合规边界与维护成本。

–

A/I shuts down - Stay human19keepitfree.ai原文 ↗

keepitfree.ai

公告宣布 A/I 服务关闭,并以 Stay human 作为收束。条目呈现的是服务生命周期结束,而不是新功能路线。

–

An Alien Mind21openai.com原文 ↗

openai.com

OpenAI 将对齐从“模型是否遵守指令”扩展到能力增强后的人类价值与治理问题,特别强调工具调用、社交交互和非语言推理会削弱仅监控 chain-of-thought 的策略。文章称 GPT-6 Astra 比 GPT-5.6 Sol 更符合当前对齐目标,但仍需要激活监测、纵深防御和国际协作来处理递归自我改进风险。

–

We monitor internal coding agents for misalignment22openai.com原文 ↗

openai.com

OpenAI 对 GPT-5.4 Thinking 的完整对话、思维链和工具调用做监测,警报目标是在 30 分钟内复核;五个月覆盖数千万条轨迹,约 1000 条中等警报,未发现最高严重度的真实对话。报告也承认低于 0.1% 的流量仍在覆盖范围外,并计划引入同步阻断,因此检测覆盖率本身是安全指标。

–

Research acceleration: The view inside OpenAI23openai.com原文 ↗

openai.com

OpenAI 披露到 8 月中旬,研究者平均每天使用超过 600 美元推理额度,90 分位超过 7000 美元;每个工作日约产生 3.1 个 agent-workday,但成功完成的 4 - 8 小时任务中逾半仍需要人工介入。7 月 20 日基础设施遭入侵后 RL 暂停两周,8 月 Astra GPU 配额下降 59.2%,其他模型只补回约 85%,显示研究加速同时受安全和算力调度约束。

–

Asahi Linux on M324asahilinux.org原文 ↗

asahilinux.org

Asahi Linux 更新 Apple Silicon M3 的移植进展,工作重点是硬件逆向、驱动和内核支持,而不是简单更换设备树。M3 的新 GPU、显示和电源管理路径需要逐项验证,意味着 Linux 可用性取决于长期上游化而非一次镜像发布。

–

NetBSD 9.5 released and EOL for NetBSD-925blog.netbsd.org原文 ↗

blog.netbsd.org

公告宣布 NetBSD 9.5 发布,并宣布 NetBSD-9 分支进入 EOL。对使用该维护线的系统而言,版本发布与维护结束需要同时纳入升级安排。

–

America's two largest school districts impose AI moratoriums27techpolicy.press原文 ↗

techpolicy.press

纽约学区禁止 K - 8 学生使用面向学生的 AI,高中只能使用批准工具;洛杉矶学区则在设备上实施一年生成式 AI 暂缓,涉及约 37.8 万名学生。政策反映出采购、隐私和教学质量尚未形成共识,学校正在用暂停争取制定边界的时间。

–

Isar Aerospace reaches orbit and deploys payloads on second flight28isaraerospace.com原文 ↗

isaraerospace.com

Isar Aerospace 2026 年 9 月 5 日从 Andøya 发射的第二次飞行完成入轨并部署 DLR/ESA Boost 载荷,经历 MaxQ、主发动机关闭、级间分离、整流罩分离和轨道圆化。它成为首家达到该里程碑的欧洲商业公司,下一步取决于量产能力和 Nova Scotia 发射场建设,而不只是单次成功。

–

The car industry A/B tested selling a car with and without CarPlay29a.wholelottanothing.org原文 ↗

a.wholelottanothing.org

同平台的 Chevrolet Blazer EV 与 Honda Prologue 只有后者提供 CarPlay/Android Auto,销量差距被当作近似 A/B 测试:Honda 优势从 2024 年 42.8% 增至 2025 年 73.1%,2026 年上半年达 165.5%。数字不能隔离品牌、价格和渠道因素,却显示车载软件入口已成为消费者可感知的配置变量。

–

How AI is breaking the British state30economist.com原文 ↗

economist.com

《经济学人》讨论 AI 在英国政府系统中的应用及其造成的治理问题。报道把焦点放在 AI 与公共治理的关系,而不是单一模型能力。

–

博客文章 · Blog Posts

13 项 · 博客文章

It took a year to ship WebAssembly in Anubis32anubis.techaro.lol原文 ↗

anubis.techaro.lol

文章回顾 Anubis 花费一年将 WebAssembly 支持送入生产,把发布过程本身作为工程主题。它说明一个功能从实现到稳定交付需要长期投入。

–

The revolt of the reader34bcantrill.dtrace.org原文 ↗

bcantrill.dtrace.org

文章称读者正在反制机器生成内容的泛滥:668 名开发者调查中,78% 会停止阅读疑似 LLM 写作,71% 会避开作者,98% 偏好人工瑕疵;Pangram 等工具约能识别四分之三文本。作者据此主张平台公开来源和人工参与度,让信任不再只靠读者猜测。

–

Your intellectual fly is open (2025)35bcantrill.dtrace.org原文 ↗

bcantrill.dtrace.org

文章用“衣服没整理好”比喻 LLM 文风留下的认知破绽,如表情符号堆叠、单句成段、固定转折和 em dash。它并不否定模型,而是把模型定位在头脑风暴、理解和编辑环节,最终观点与声音仍须由作者承担。

–

Doomscrolling Ourselves to Death36edwest.co.uk原文 ↗

edwest.co.uk

文章把无尽信息流与阅读能力、政治讨论和心理状态的收缩联系起来:平均每天查看手机约 144 次,TikTok 半小时约 260 个视频,并引用挪威电视覆盖与 IQ 每年下降约 0.08 的相关研究。论证的重点是注意力结构如何把复杂经验压成短促情绪,而非简单谴责某个应用。

–

The purpose of DNS is to spread scams37simonwillison.net原文 ↗

simonwillison.net

Simon Willison 转述 Interisle 数据,讨论新域名注册与诈骗滥用的关系。文章把域名注册视为诈骗传播链条的一环,标题以反讽方式强调这种用途。

–

Introducing GPT-6 Astra for developers38simonwillison.net原文 ↗

simonwillison.net

GPT-6 Astra 的演示把改进落在细节保持和多约束提示理解:从花园、船厂到戴森球的 3D 场景都能维持结构,鹈鹕骑车并佩戴红色领巾的示例则检验组合约束。对开发者而言,真正需要验证的是这类一致性在长流程和真实数据上的稳定成本。

–

How Swiss tables work in Go built-in map39victoriametrics.com原文 ↗

victoriametrics.com

文章拆解 Go 1.24 map 的 Swiss Table:控制字节、H1/H2 片段和组级探测减少无效比较;提示容量 100 会落到 128 槽、112 个可用条目,负载阈值为 87.5%。达到 1024 槽后再拆表,以及 Go 1.27 的 mapsplitgroup 实验,说明性能优化同时受内存布局和未来兼容性牵引。

–

AI, Tools and Transformation44ben-evans.com原文 ↗

ben-evans.com

Ben Evans 指出企业常有数百乃至数千个应用,AI 让定制工具更便宜,却不会自动决定哪些流程值得重做。约一半试点最终变成日常工作,另一半停留在演示,分水岭是流程重构和变更管理;软件将同时出现制度化系统与临时化工具。

–

引用来源 · References

58 条 · 引用
  1. 1 Harnessing the Universal Geometry of Embeddings. arXiv:2505.12540https://arxiv.org/abs/2505.12540 ↩ 回到正文 · back to text
  2. 2 Large-Language Models as a Cognitive Virus. arXiv:2609.03344https://arxiv.org/abs/2609.03344 ↩ 回到正文 · back to text
  3. 3 indic-modernBERThttps://github.com/kkkamur07/indic-modernBERT ↩ 回到正文 · back to text
  4. 4 Wayfinderhttps://github.com/DivakarUngatla/wayfinder ↩ 回到正文 · back to text
  5. 5 OKF Agent Memoryhttps://github.com/okf-memory/okf-agent-memory ↩ 回到正文 · back to text
  6. 6 Madorhttps://github.com/marsbos/mador ↩ 回到正文 · back to text
  7. 7 MaskShifthttps://github.com/nafeeur/MaskShift ↩ 回到正文 · back to text
  8. 8 Byosynchhttps://byosynch.com/ ↩ 回到正文 · back to text
  9. 9 Agentic OShttps://github.com/mmeyerlein/meclaw ↩ 回到正文 · back to text
  10. 10 toologhttps://github.com/zaghaghi/toolog ↩ 回到正文 · back to text
  11. 11 ZeroThesishttps://zerothesis.com ↩ 回到正文 · back to text
  12. 12 HakoVMhttps://github.com/las7/HakoVM ↩ 回到正文 · back to text
  13. 13 Twinrunhttps://github.com/prian3003/twinrun ↩ 回到正文 · back to text
  14. 14 dsnitchhttps://github.com/infomaniac777/dsnitch ↩ 回到正文 · back to text
  15. 15 vernLLMhttps://github.com/LakBud/vernLLM ↩ 回到正文 · back to text
  16. 16 webmcp-stackhttps://github.com/SouravInsights/webmcp-stack ↩ 回到正文 · back to text
  17. 17 XHFShttps://codeberg.org/futureg-lab/xhfs ↩ 回到正文 · back to text
  18. 18 Following legal advice, the Nitter project will continuehttps://github.com/zedeus/nitter ↩ 回到正文 · back to text
  19. 19 A/I shuts down - Stay humanhttps://keepitfree.ai/announcements/a/i-shuts-down-stay-human/ ↩ 回到正文 · back to text
  20. 20 QBittorrent breaks out of sandbox to commit crimeshttps://beige.party/@intransitivelie/117057396732763183 ↩ 回到正文 · back to text
  21. 21 An Alien Mindhttps://openai.com/index/an-alien-mind/ ↩ 回到正文 · back to text
  22. 22 We monitor internal coding agents for misalignmenthttps://openai.com/index/how-we-monitor-internal-coding-agents-misalignment/ ↩ 回到正文 · back to text
  23. 23 Research acceleration: The view inside OpenAIhttps://openai.com/index/research-acceleration-view-inside-openai ↩ 回到正文 · back to text
  24. 24 Asahi Linux on M3https://asahilinux.org/2026/09/m2-episode-1/ ↩ 回到正文 · back to text
  25. 25 NetBSD 9.5 released and EOL for NetBSD-9https://blog.netbsd.org/tnf/entry/netbsd_9_5_released_and ↩ 回到正文 · back to text
  26. 26 Chrome again exempts Google from user site data settingshttps://lapcatsoftware.com/articles/2026/9/1.html ↩ 回到正文 · back to text
  27. 27 America's two largest school districts impose AI moratoriumshttps://www.techpolicy.press/americas-two-largest-school-districts-impose-ai-moratoriums/ ↩ 回到正文 · back to text
  28. 28 Isar Aerospace reaches orbit and deploys payloads on second flighthttps://isaraerospace.com/press/history-for-european-spaceflight-isar-aerospace-reaches-orbit-and-deploys-payloads-on-second-flight ↩ 回到正文 · back to text
  29. 29 The car industry A/B tested selling a car with and without CarPlayhttps://a.wholelottanothing.org/the-car-industry-a-b-tested-selling-the-same-car-with-and-without-carplay-and-the-results-are-not-shocking/ ↩ 回到正文 · back to text
  30. 30 How AI is breaking the British statehttps://www.economist.com/leaders/2026/08/06/how-ai-is-breaking-the-british-state ↩ 回到正文 · back to text
  31. 31 Cloud in a Bottle: making self-hosting accessible to everyonehttps://cloudinabottle.org/blog/launch-post ↩ 回到正文 · back to text
  32. 32 It took a year to ship WebAssembly in Anubishttps://anubis.techaro.lol/blog/2026/anubis-wasm/ ↩ 回到正文 · back to text
  33. 33 Have the frontier labs mixed up AI safety and security?https://martinalderson.com/posts/ai-safety-vs-security/ ↩ 回到正文 · back to text
  34. 34 The revolt of the readerhttps://bcantrill.dtrace.org/2026/09/05/the-revolt-of-the-reader/ ↩ 回到正文 · back to text
  35. 35 Your intellectual fly is open (2025)https://bcantrill.dtrace.org/2025/12/05/your-intellectual-fly-is-open/ ↩ 回到正文 · back to text
  36. 36 Doomscrolling Ourselves to Deathhttps://www.edwest.co.uk/p/doomscrolling-ourselves-to-death ↩ 回到正文 · back to text
  37. 37 The purpose of DNS is to spread scamshttps://simonwillison.net/2026/Sep/6/the-purpose-of-dns-is-to-spread-scams/ ↩ 回到正文 · back to text
  38. 38 Introducing GPT-6 Astra for developershttps://simonwillison.net/2026/Sep/5/introducing-gpt-6-astra-for-developers/ ↩ 回到正文 · back to text
  39. 39 How Swiss tables work in Go built-in maphttps://victoriametrics.com/blog/go-swiss-table-map/index.html ↩ 回到正文 · back to text
  40. 40 Terence Tao on “prematurely solving [a maths] problem by purely AI-powered methods”https://mathstodon.xyz/@tao/117207856734787448 ↩ 回到正文 · back to text
  41. 41 There's No Limit to How Bad Code Can Gethttps://zachkehs.com/blog/theres_no_limit_to_how_bad_code_can_get/ ↩ 回到正文 · back to text
  42. 42 Controlling when CSS custom property values are computedhttps://jakearchibald.com/2026/css-custom-property-compute-time/ ↩ 回到正文 · back to text
  43. 43 Recreating Minecraft Is Not a Benchmarkhttps://kuber.studio/blog/Reflections/Recreating-Minecraft-is-Not-a-Benchmark ↩ 回到正文 · back to text
  44. 44 AI, Tools and Transformationhttps://www.ben-evans.com/benedictevans/2026/9/3/ai-tools-and-transformation ↩ 回到正文 · back to text
  45. 45 tailscale/tailcathttps://github.com/tailscale/tailcat ↩ 回到正文 · back to text
  46. 46 magnitudedev/magnitudehttps://github.com/magnitudedev/magnitude ↩ 回到正文 · back to text
  47. 47 checkstyle/checkstylehttps://github.com/checkstyle/checkstyle ↩ 回到正文 · back to text
  48. 48 BraveOPotato/FckSignupshttps://github.com/BraveOPotato/FckSignups ↩ 回到正文 · back to text
  49. 49 k2-fsa/OmniVoicehttps://github.com/k2-fsa/OmniVoice ↩ 回到正文 · back to text
  50. 50 actions/upload-artifacthttps://github.com/actions/upload-artifact ↩ 回到正文 · back to text
  51. 51 huggingface/datasetshttps://github.com/huggingface/datasets ↩ 回到正文 · back to text
  52. 52 thewh1teagle/vibehttps://github.com/thewh1teagle/vibe ↩ 回到正文 · back to text
  53. 53 cobusgreyling/loop-engineeringhttps://github.com/cobusgreyling/loop-engineering ↩ 回到正文 · back to text
  54. 54 code-yeongyu/lazycodexhttps://github.com/code-yeongyu/lazycodex ↩ 回到正文 · back to text
  55. 55 Jakubantalik/Libraries.devhttps://github.com/Jakubantalik/Libraries.dev ↩ 回到正文 · back to text
  56. 56 humanlayer/skillshttps://github.com/humanlayer/skills ↩ 回到正文 · back to text
  57. 57 WorldFlowAI/everything-claude-codehttps://github.com/WorldFlowAI/everything-claude-code ↩ 回到正文 · back to text
  58. 58 coulsontl/ai-toolboxhttps://github.com/coulsontl/ai-toolbox ↩ 回到正文 · back to text