每日 Harness 开源 · Source
全部刊期 · All issues

每日 Harness

2026-09-08 · Tuesday, September 8, 2026

智能体走向持久可控与实证验证

视图 · View

今日重点 · Today's Highlights

论文 · Papers

15 项 · 论文

A Removal Based Approach to Improve LLM Faithfulness at Test-Time6arxiv.org原文 ↗

arxiv.org

论文把解释不忠实拆成遗漏真实影响的“不完整”和捏造影响的“不健全”,在推理时删除解释没有提及的输入概念后重新提问。两个数据集、多个模型家族和两种独立指标均显示其忠实度超过普通提示及“请保持忠实”提示;无需权重访问,代价是需要额外一次查询和概念归因步骤。

–

What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents7arxiv.org原文 ↗

arxiv.org

研究固定 Aider、OpenHands、Qwen Code、SWE-agent 的任务记录,只比较 GRPO 在同一任务内按 harness 分组或跨 harness 混组。24,000 次密封 SWE-bench 评估里,执行 harness 把平均解决率从 2.14% 推到 9.27%,而分组规则在留出 harness 上仅 +0.25 个百分点且置信区间跨零,信用分配并非主要瓶颈。

–

MaxKernel: Agentic Kernel Generation for TPUs8arxiv.org原文 ↗

arxiv.org

MaxKernel 用规划、实现、自调试、测试和 profiling 子代理串起三种 TPU kernel 搜索:人机协作、自动指标循环和图式全局探索。它在 50 项 JaxBench 任务及真实开源模型 workload 上达到专家手调实现的性能水平;论文的实际贡献在于把编译器反馈变成多代理共同的搜索信号,而不只是让模型一次性写 kernel。

–

SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents9arxiv.org原文 ↗

arxiv.org

SiLR 把违规后的准入门建模为轨迹搜索,shadow-execute 每个提案,并按“有多少分支仍过载、各分支严重度”组成的乘积序保留结构信息。作者证明不存在对该序关系声音的标量替代;Gym-ANM 的多动作 episode 取得 21/21 恢复,对比终止式 0/21 和最佳标量门 9/21,说明阈值调参无法修复表示缺陷。

–

本期重点From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents4arxiv.org原文 ↗

arxiv.org

这项工作把每轮 GUI 轨迹与评估器反馈写进可审计技能库,更新只在下一轮快照生效,因此能测量技能的增量而不是把能力提升归因于模型微调。四个 OSWorld 域在五轮空库热身后提升 5.7 - 18.6 个百分点;GIMP 中跨任务检索和 revision churn 也表明,重复接受编辑不等于找回最初任务。

–

Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing10arxiv.org原文 ↗

arxiv.org

RSM-full 以余弦门控的 max-member merge 控制写入,以 atom-aware packer 组织读取,把“保留哪些记忆”和“如何装入窗口”作为一体化设计。AMA-Bench 的 4k 预算达到完整上下文 83% 质量、token 成本 32%;在约 2.6k - 5k 区间比 Online K-Means 高 3.5 - 6.0 个百分点(四种子,p<.001),是针对紧凑预算的实测 Pareto 点。

–

TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing11arxiv.org原文 ↗

arxiv.org

TROVE 从已评估 workflow-search 轨迹提取复合技能和 outcome-conditioned 图,在线只编辑被新证据推翻的路线后缀,避免整条计划重算。代码、问答、数学三类任务都优于数据集级优化和图调度基线;消融把收益拆成复合技能的离线提升、插入动作的局部纠错,以及后缀替换带来的计算节省。

–

From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments12arxiv.org原文 ↗

arxiv.org

综述按委托权限、时间持久性和环境耦合重新整理代理系统,并把模型、harness、环境三层分开,避免把接口扩张误写成自主性。基于截至 2026-08-31 的原始研究和官方规格,作者判断动作接口的证据最充分,而稳健完成、恢复、授权和独立验证仍没有同等强度的支持;这是一条对“能力叙事”降温的证据框架。

–

Substrate-Aware AI Agents: Execution Context as a First-Class Input13arxiv.org原文 ↗

arxiv.org

研究把内存、运行时间、算力和操作约束显式写进规划状态,以高维两两距离代码任务检验不再“substrate blind”。给出 128 MB RAM、10 秒 wall-time 合约后,14 个对齐比较中 13 个峰值内存下降,三组模型平均耗时全部下降,最快达到 3.1 倍;约束信息改变的是实现选择,不只是最终筛选。

–

Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability14arxiv.org原文 ↗

arxiv.org

研究将同一历史分别保存为原文、RAG、自然语言 NOTES 和固定 schema 知识图,测试写入模型更换后的读取、检索与修复。48 条随机答案码历史显示 KG-fixed 的准确率变化为 +0.0004±0.0020,而 NOTES 会随迁移方向出现 +9.91 或 −13.28 个百分点;记忆格式比“保留同一数据库”更决定升级风险。

–

CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents15arxiv.org原文 ↗

arxiv.org

App-Forge、Task-Weave、Path-Steer 三段式管线把真实桌面应用转成共享状态的 GUI+CLI 环境,并把高效、验证过的混合路径用于后训练。9B 模型在 CUA-Verse 得分提升 39.3、步骤减少 37%、token 减少 60%;迁移到 OSWorld 仍有成功率 +16.8 个百分点和步骤 −57%,说明收益来自接口协同而非单纯视觉增强。

–

RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents16arxiv.org原文 ↗

arxiv.org

RefactorPlatform 将模型、执行模式和提示具体度拆成可控轴,在隔离 workspace 中记录 diff、token、transcript,并用 AST 校验跨文件行为是否保持。100 个 RefactorBench 任务中,AST 感知分块比 token 窗口高 25 - 30%;精简检索单代理通过率 86% 胜过评测中的子代理 66%,检索的 token 开销被准确率收益抵消后,单次成功成本不变。

–

本期重点CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls2arxiv.org原文 ↗

arxiv.org

CONTINUITY 用签名根授权、provenance commitment、角色绑定转移收据、类型化释放和 effect-bound permit,把每个组件的假设与保证连成端到端授权见证。参考实现覆盖 32 类故障;2,560 个参数化攻击实例无一提交有害外部效果,同时完成 700 个良性任务并升级 200 个含糊任务,重点在跨边界不丢失语义而非增加单点过滤器。

–

本期重点Agentic Context Cracking: Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data3arxiv.org原文 ↗

arxiv.org

cracking 子代理从已加载文档上下文分叉,按观测到的问题决定抽取哪些结构,并对相关未来查询做投机性整理。论文指出在 FanOutQA 上理想结构化存储可便宜 28 倍;只扩展一个相关问题时降本 53% 且准确率保持,代价是要维护结构化结果的依据链和更新策略。

–

本期重点Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation1arxiv.org原文 ↗

arxiv.org

适配器先通过代码审查和 parity experiment 把 80 多个 benchmark 统一进 Harbor,再用 54 个 benchmark、8 个模型分析能力与失败模式,最后过滤成 29 个 benchmark 的 82 个高质量任务。所有模型-harness 配置通过率低于 30%,最强 GPT-5.5+Codex 为 28.0%;这种刻意保持的难度为跨项目比较留下余量,也暴露出当前代理离稳健完成仍远。

–

开源 / 项目 · Projects

15 项 · 开源 / 项目

本期重点Keyclasp - Let agents use tokens without putting them in prompts5github.com原文 ↗

github.com

Keyclasp 在本地加密 vault 保存凭据,命令执行前按名称注入实际 token,避免 secret 出现在提示词、代码文件和常规输出里。README 的安全边界很具体:被注入的子进程仍能把值发往网络、写入磁盘或交给同用户进程,因此它解决的是上下文暴露,不是沙箱问题。

–

Worldfixture - An OSS fake company your software can interact with17worldfixture.com原文 ↗

worldfixture.com

Worldfixture 用本地可运行的假公司模拟支付、邮件、存储和协作服务,让应用在无真实 SaaS 账号的情况下联调业务流程。项目把“外部依赖的替身”做成共享测试面,适合验证 agent 的多服务调用与错误处理,而不是只测单个 API mock。

–

Expose - a self-hosted tunneling tool like ngrok/localtunnel18github.com原文 ↗

github.com

Expose 是 Go 单二进制隧道,把 localhost 暴露为可分享 URL,主打 webhook 测试、移动调试和现场演示。零配置、免注册和自托管降低了把临时服务交给第三方的门槛;仓库同时保留 internal、文档和 GoReleaser,便于审查部署链。

–

MobileCode - OpenCode with React Native Previews19github.com原文 ↗

github.com

MobileCode 从 OpenCode 分叉,补上 iOS/Android 构建与 React Native 预览,让 agent 的编辑循环能看到真实移动端产物。它的价值不在新模型,而在把编译、运行和视觉回馈纳入 OpenCode 工作流,减少“代码看似完成但设备行为未知”的盲区。

–

OpenLivery - open-source multi-tenant WhatsApp agents for agencies20github.com原文 ↗

github.com

OpenLivery 提供代理机构管理客户 agent 的多租户 workspace、品牌门户、WhatsApp/网页聊天、知识库、自定义工具和人工转接。文档将 Docker 快速启动、租户隔离、密钥配置、备份升级分别列出,并允许自带 OpenAI/Anthropic key;它更像可运营的白标后台,而非单一聊天 demo。

–

Crew - Let Claude/Codex/OpenCode agents talk to each other21github.com原文 ↗

github.com

Crew 把三种 CLI 会话的状态、摘要和 transcript 尾部自动注入彼此上下文,并提供数秒内送达的 `crew send` agent-to-agent 邮件。全局 npm 安装会自动写入 Claude、Codex、OpenCode hook,支持 JSON 监控;它选择共享 checkout,因此并行修改的协调责任仍在用户和 agent。

–

Benzi - Code Intelligence Infrastructure for Frontier AI Models22github.com原文 ↗

github.com

Benzi 让 tree-sitter 编译器先解析每个文件、解析 import/继承/调用/引用,再从统一 query map 回答符号问题,而不是把仓库全文塞给模型。作者称十种语言共享同一图结构,并把调用流和数据流在调用点合并;这种 O(1) 查询定位的是跨文件事实,不能替代对行为语义的测试。

–

Coop - Isolated VM Environments for Running Claude Code and Codex23github.com原文 ↗

github.com

Coop 以 Rust CLI 管理 disposable VM,让 Claude Code/Codex 在获得 Docker、编译器和包管理器权限时与宿主机隔离。每个环境可重复创建和销毁,适合运行未知依赖、破坏性测试和无人值守任务;隔离粒度由 VM 配置决定,不能被描述成自动的应用级安全证明。

–

Engrim - A universal, local-first SQLite memory engine for AI CLIs24github.com原文 ↗

github.com

Engrim 把项目决策、约束和工作状态浓缩为约 4,000 字符的 episodic memory,客户端可经 MCP 自动写入,也能用 `engrim add` 手动保存。它把记忆从模型供应商解耦,支持 Antigravity、Claude Code、Cursor、Windsurf、Codex 之间切换;核心取舍是策展和检索质量,而非无限堆积聊天记录。

–

Chrome-bridge - let any AI agent drive your real logged-in Chrome25github.com原文 ↗

github.com

Chrome-bridge 通过一个扩展和零依赖 Node CLI 操作用户现有的登录 tab、SSO 和 profile,不要求 MCP-capable 客户端或重新启动浏览器。它提供 a11y snapshot、元素引用、CDP 截图和网络捕获,Chrome 117+/Node 18+ 即可运行;便利性同时意味着 agent 与真实账户权限处于同一浏览器边界。

–

GuidedReview - Review AI-generated code before you sign your name to it26github.com原文 ↗

github.com

GuidedReview 将本地 diff 和 GitHub PR 聚类成有顺序的 review units,LLM 只负责把 hunks 按 schema、逻辑、调用点、测试组织并加短评。CLI 与 Chrome 扩展直连用户模型和 GitHub,没有后端代审;代码来源始终是真实 diff,人的逐段确认仍是流程中的决策点。

–

DashClaw - policy and approval layer for unattended coding agents27github.com原文 ↗

github.com

DashClaw 把无人值守 agent 的远程审批、策略检查和执行证据集中在一层,并配套 SDK、MCP 配置、Dockerfile 与 smoke tests。它补足的是“动作能否执行、谁批准、留下什么证据”的治理链,而不是再造一个代码生成器;高风险策略的完整性仍取决于接入方是否覆盖所有执行入口。

–

A local visual tool cli/mcp for agents to propose architecture changes28github.com原文 ↗

github.com

WorkBraid 让人和 agent 在本地工作台中维护组件图、提交架构 proposal、比较版本并导出 PDF 报告,CLI 与 MCP 共用运行中的服务。源码运行要求 Go 1.26+、Node 24+,`workbraid mcp` 连接已有 UI;它把架构讨论从聊天文本变成可审阅的持久对象。

–

Awareness-Market - Open-source agent memory layer, 96% on LongMemEval, local-first29github.com原文 ↗

github.com

Awareness Local 以 Markdown 记忆卡为持久层,用 SQLite FTS5 与 embedding 混合检索,并通过 MCP 服务 Cursor、Claude Code、Copilot、Cline。一次 `npx @awareness.market/setup` 即可离线启动,还提供决策/风险知识卡和 Web dashboard;可解释的文本存储让人工纠正比黑盒向量库更直接。

–

MathKernel: An evidence-aware multi-engine mathematics kernel and MCP server30github.com原文 ↗

github.com

MathKernel 让 LLM 负责解析和解释,把精确、符号、形式化、区间证书和数值计算交给带 typed MathIR 的内核,并为每个结论记录 provenance 与 trust label。README 特别提醒“引擎一致不等于证明”,要求 claim-specific evidence bundle;MCP server、Python 库和便携产物把同一证据链暴露给不同工作流。

–

行业动态 · Industry News

14 项 · 行业动态

Tesla killing Solar Roof is leaving installers with six-figure losses32electrek.co原文 ↗

electrek.co

Tesla 8 月停止供应 Solar Roof,转向传统光伏板,约七年在美国只完成 3,000 套、峰值每周 21 - 32 套,远低于每周 1,000 套承诺。安装商为多周培训、专用工具和施工团队投入数十万至数十万美元,如今还要面对未完工项目、零件和 25 年保修的不确定性;撤掉产品线的成本被转移给了合作网络。

–

This Month in Ladybird - August 202633ladybird.org原文 ↗

ladybird.org

Ladybird 的 2026 年 8 月 newsletter 汇总浏览器项目月度开发进展,条目本身没有额外的技术数字。它的读法应是持续观察独立浏览器在引擎、平台和工程基础设施上的累积变化,而非把月报标题当成单次发布事件。

–

Scientists observe Einstein's gravity in the quantum world34ox.ac.uk原文 ↗

ox.ac.uk

实验把超冷铷原子波分成一条被磁场托住、另一条自由下落的路径,再合并测量量子相位;测得相位与将等效原理应用于量子波的预测一致。牛津明确划出结论边界:这是首次直接测量自由落体的预测量子相位,不是证明引力本身量子化,也未达到检验宏观叠加崩溃所需质量和时长。

–

Smartphone makers don't bother to comply with EU repairability requirements35theregister.com原文 ↗

theregister.com

The Register 指出手机厂商对欧盟可维修性要求的执行并不充分,争议落在法规指标与实际拆修体验之间的落差。报道的意义不只在合规名单,也在于维修评分、零件供应和软件支持若不能联动,消费者看到的“可修”标签并不等于可持续使用年限。

–

Switzerland's Federal Government Is Replacing Microsoft on 3k Computers36itsfoss.com原文 ↗

itsfoss.com

瑞士联邦政府以 CHF 900 万启动 3,000 台工作站的 openDesk 试点,目标 2027 年底完成,期间与 Microsoft 365 并行;此前 172 人 PoC 对文档和邮件满意,但视频会议仍有技术限制。军方 Cyber Command 计划 2026 年 10 月前全量迁移,背后是外国访问风险、供应商依赖和授权成本,属于以渐进迁移换取数字主权的工程案例。

–

Research acceleration: The view inside OpenAI37openai.com原文 ↗

openai.com

OpenAI 披露中位研究员到 8 月中每天消耗超过 600 美元 API 推理,90 分位超过 7,000 美元,组织总体达到每个人类工作日 3.1 个 agent-workday;但过去六个月成功的 4 - 8 小时任务中超过一半仍有人工介入。安全限制使 Astra GPU 配额再降 59.2%,其他模型配额增加 17.2% 抵消约 85%,数据同时显示自动化扩张与算力/安全治理瓶颈并存。

–

PostgreSQL 19 Interactive Tour38victoriametrics.com原文 ↗

victoriametrics.com

VictoriaMetrics 用交互式示例导览 PostgreSQL 19 的新特性,重点是让读者在可执行上下文中理解行为变化,而非只读 release note。该条适合作为版本升级前的动手入口,具体 SQL 和执行器差异应结合官方文档与自己的 workload 验证。

–

OpenAI brings back 5 hour limit for plus and business standard users39news.ycombinator.com原文 ↗

news.ycombinator.com

Hacker News 讨论 OpenAI Plus 与 Business Standard 恢复五小时使用时长限制,帖子体现的是用户对配额政策变化的即时反馈,而非完整官方公告。对依赖长会话或并发 agent 的团队来说,限制的时间窗和套餐边界比“恢复限制”四个字更需要核对。

–

VMware migration reduces Tottenham Hotspur's licensing fees by 85 percent40arstechnica.com原文 ↗

arstechnica.com

热刺在三个月内从 VMware 迁到 HPE Morpheus VME/OpsRamp,授权费降幅超过 85%,但 CTO 强调把虚拟化、监控和 AI 运维放进统一界面才是主要收益。约 35 人的技术团队借自动发现和修复故障来放大人力;案例说明迁移回报同时由许可账单和运维编排能力决定。

–

Initial effects of AI technology on employment look positive41economist.com原文 ↗

economist.com

The Economist 汇总的早期研究把 AI 对就业的初始影响描述为偏正面,标题所说的“jobs apocalypse postponed”强调观察窗口仍短。解读这类宏观结论需区分新增岗位、任务重分配和生产率收益,不能把短期净变化直接外推到长期职业结构。

–

Nvidia's Jensen Huang says 'AGI has arrived' and congratulates OpenAI42businessinsider.com原文 ↗

businessinsider.com

Jensen Huang 公开表示“AGI 已到来”并祝贺 OpenAI,这是一位产业领袖的判断性表态,不是带评测协议和可复现实验的科学结论。把它放入产业语境更合适:定义边界和证据标准仍未因一句庆祝性发言而统一。

–

TiVo to charge money for skipping commercials in your own recordings44cordcuttersnews.com原文 ↗

cordcuttersnews.com

TiVo 从 2026-11-02 起取消免费 SkipMode,改以 30 天试用后收费的 Premium Auto Commercial Skip 提供自动/一键跳过,价格尚未公布;旧录制的标记仍会工作,手动 30 秒跳过和快进不受影响。把曾经的核心 DVR 差异化改成订阅项,反映 Xperi 已从硬件制造转向电视系统和广告技术。

–

博客文章 · Blog Posts

15 项 · 博客文章

The Education of a Doomer45borretti.me原文 ↗

borretti.me

这篇随笔回看作者的技术观念与悲观主义如何形成,重点在思想路径而非某个工程结论。它适合与“技术进步必然向好/必然失控”的单线叙事对照阅读,保留个人经验的局限也正是文本的一部分。

–

Decapitating a MacBook (2025)46mm-dev.rocks原文 ↗

mm-dev.rocks

作者为 Apple 平台开发买下屏幕损坏的 8GB M1 MacBook Air(约 £300),拆屏改成 headless 机器,绕开 macOS 虚拟机性能和云 CI 费用问题。文章逐段处理 USB-C 显示、供电、散热和远程访问,最终证明低功耗二手硬件也能支撑 Flutter 测试;价值在于记录真实摩擦而非展示一次性改装照片。

–

Rebuilding a 1995 GPS Time Server so I don't get Telstra'd47jeffgeerling.com原文 ↗

jeffgeerling.com

Geerling 用 Pi 5、GNSS HAT、gpsd 和 Chrony 把 1995 年 TrueTime XL-AK 恢复为 stratum-1 NTP 服务器,并复用原机 LCD/LED。固定风扇、force_turbo、隔离 PPS 中断和给晶振保温后,他在 6 - 12 小时测试中看到更稳定的频率,持续功耗约多 1W;这是一份把硬件历史转成可运行基础设施的实作笔记。

–

De-Brainrot Vacations48devz.cl原文 ↗

devz.cl

作者把假期设为减少数字媒体消费的自我实验,观察注意力和日常节奏如何变化。文章的技术性不在工具,而在把“少刷内容”当成可观察的行为干预;结论应理解为个人样本,不能直接当作普遍生产率方案。

–

Making a Python interpreter in 1024 bytes49austinhenley.com原文 ↗

austinhenley.com

文章挑战在 1024 字节内实现可运行的 Python 解释器,把语法、运行时和压缩策略同时推到极限。这个尺度迫使作者明确“语言实现最小闭环”需要哪些部件,也让字节级取舍比常规性能优化更可见;适合作为解释器结构的反向教学案例。

–

How well do agents use test/verification techniques?50danluu.com原文 ↗

danluu.com

Dan Luu 观察编码代理在真实流程中如何编写、运行和解释测试,问题焦点是验证行为是否跟得上生成速度。条目提醒读者区分“代理写出了测试”与“测试真的约束了结果”;评价 agent 时应看它是否主动复现失败、缩小假设并检查边界,而非只数测试文件。

–

Rust debugging survey 2026 results51blog.rust-lang.org原文 ↗

blog.rust-lang.org

Rust 官方发布 2026 调试体验调查结果,面向工具链和开发反馈而非语言新特性。其价值在于把编译器错误、调试器和运行时诊断作为一个完整工作流来衡量,后续应结合原始问卷数据解读痛点排序。

–

Programming is Art52orchidfiles.com原文 ↗

orchidfiles.com

文章讨论编程实践与艺术创作的相似处,关注选择、约束、表达和迭代如何共同形成作品。它并不提供工程指标,却能解释为什么代码质量常包含不可量化的结构感、节奏和取舍。

–

C Is Not a Low-Level Language (2018)53queue.acm.org原文 ↗

queue.acm.org

ACM Queue 文章重新追问 C 的“低级”标签,把抽象层次从语法表面移到编译器、内存模型和硬件行为的组合。这个视角提醒系统开发者:能写指针不等于直接控制机器,C 同样依赖大量未显式写出的运行时与平台约定。

–

Bot Detection Without JavaScript: What My Blog Measured54gkoreli.com原文 ↗

gkoreli.com

作者用 Cloudflare Worker/D1 记录 ASN、Fetch Metadata、Accept 头和 UA,强调“客户端是谁”与“是否有人阅读”不是同一问题。七天窗口里浏览器 UA 有 1,209 次页面事件、578 个日标识,Cloudflare 仪表盘只有 113 次页面、52 次访问;一个移动标识一秒抓 31 页,促使他改用统一时间窗、事件边界和可复核规则,而非寻找神奇比例。

–

Demystifying complex configurations55guix.gnu.org原文 ↗

guix.gnu.org

Guix 博文从组织和调试角度解释复杂配置,目标是让声明式系统的依赖关系和失败位置更可见。它对应的工程问题不是“少写几行配置”,而是建立能被拆解、复现和局部验证的配置结构。

–

Data races and the limits of ThreadSanitizer in C and Go56theconsensus.dev原文 ↗

theconsensus.dev

文章讨论 ThreadSanitizer 在 C 与 Go 数据竞争检测中的覆盖边界,提醒动态工具只能观察被执行到的路径和满足其插桩假设的内存访问。工程上仍需把静态分析、锁设计、模型检查和代码审查组合起来,不能把“TSan 没报错”当作无竞争证明。

–

Visualizing Rust's Vtables: How dyn Trait Works In Memory57sofiabelen.github.io原文 ↗

sofiabelen.github.io

文章先用 C++ 虚函数和 CRTP 建立对照,再解释 Rust 泛型的静态分发、`dyn Trait` 的 vtable 指针,以及每个 (Type, Trait) 对应一张表。它还说明 object safety 为何禁止返回 `Self` 或带泛型参数的方法;把抽象规则落到内存布局,能避免把 Rust 当成“换语法的 C++”。

–

.gitignore Everything by Default58packagemain.tech原文 ↗

packagemain.tech

作者建议用 `*` 默认忽略,再用 `!` 显式放行源码、README、go.mod 等文件,以免把 node_modules、IDE 文件和环境变量带进仓库。示例同时承认这不是普适规范,并给出 `git check-ignore -v` 的排查方式;它把版本控制从“事后清理”改成“允许清单”决策。

–

There's No Limit to How Bad Code Can Get59simonwillison.net原文 ↗

simonwillison.net

Simon Willison 认为直接重写技术债务系统很少成功:旧系统仍是移动靶,新团队又无法完整理解行为,最后常留下两个生产系统。更稳妥的路径是先补自动化测试,再通过小步迁移和针对性重构交付价值;这与等待“大爆炸”替换形成鲜明的风险差异。

–

引用来源 · References

69 条 · 引用
  1. 1 Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation. arXiv:2609.04298https://arxiv.org/abs/2609.04298 ↩ 回到正文 · back to text
  2. 2 CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls. arXiv:2609.05269https://arxiv.org/abs/2609.05269 ↩ 回到正文 · back to text
  3. 3 Agentic Context Cracking: Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data. arXiv:2608.31082https://arxiv.org/abs/2608.31082 ↩ 回到正文 · back to text
  4. 4 From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents. arXiv:2609.04869https://arxiv.org/abs/2609.04869 ↩ 回到正文 · back to text
  5. 5 Keyclasp - Let agents use tokens without putting them in promptshttps://github.com/AndreaCatalucci/keyclasp ↩ 回到正文 · back to text
  6. 6 A Removal Based Approach to Improve LLM Faithfulness at Test-Time. arXiv:2609.04343https://arxiv.org/abs/2609.04343 ↩ 回到正文 · back to text
  7. 7 What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents. arXiv:2609.04518https://arxiv.org/abs/2609.04518 ↩ 回到正文 · back to text
  8. 8 MaxKernel: Agentic Kernel Generation for TPUs. arXiv:2609.04523https://arxiv.org/abs/2609.04523 ↩ 回到正文 · back to text
  9. 9 SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents. arXiv:2609.04629https://arxiv.org/abs/2609.04629 ↩ 回到正文 · back to text
  10. 10 Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing. arXiv:2609.04915https://arxiv.org/abs/2609.04915 ↩ 回到正文 · back to text
  11. 11 TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing. arXiv:2609.05019https://arxiv.org/abs/2609.05019 ↩ 回到正文 · back to text
  12. 12 From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments. arXiv:2609.04894https://arxiv.org/abs/2609.04894 ↩ 回到正文 · back to text
  13. 13 Substrate-Aware AI Agents: Execution Context as a First-Class Input. arXiv:2609.05232https://arxiv.org/abs/2609.05232 ↩ 回到正文 · back to text
  14. 14 Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability. arXiv:2609.05339https://arxiv.org/abs/2609.05339 ↩ 回到正文 · back to text
  15. 15 CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents. arXiv:2609.05374https://arxiv.org/abs/2609.05374 ↩ 回到正文 · back to text
  16. 16 RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents. arXiv:2609.04898https://arxiv.org/abs/2609.04898 ↩ 回到正文 · back to text
  17. 17 Worldfixture - An OSS fake company your software can interact withhttps://worldfixture.com/ ↩ 回到正文 · back to text
  18. 18 Expose - a self-hosted tunneling tool like ngrok/localtunnelhttps://github.com/kernelshard/expose ↩ 回到正文 · back to text
  19. 19 MobileCode - OpenCode with React Native Previewshttps://github.com/hsandhu/mobilecode ↩ 回到正文 · back to text
  20. 20 OpenLivery - open-source multi-tenant WhatsApp agents for agencieshttps://github.com/sarrazola/openlivery ↩ 回到正文 · back to text
  21. 21 Crew - Let Claude/Codex/OpenCode agents talk to each otherhttps://github.com/0xmmo/crew/ ↩ 回到正文 · back to text
  22. 22 Benzi - Code Intelligence Infrastructure for Frontier AI Modelshttps://github.com/oooscoos/Benzi ↩ 回到正文 · back to text
  23. 23 Coop - Isolated VM Environments for Running Claude Code and Codexhttps://github.com/trailofbits/coop ↩ 回到正文 · back to text
  24. 24 Engrim - A universal, local-first SQLite memory engine for AI CLIshttps://github.com/timgordontg/engrim ↩ 回到正文 · back to text
  25. 25 Chrome-bridge - let any AI agent drive your real logged-in Chromehttps://github.com/siropkin/chrome-bridge ↩ 回到正文 · back to text
  26. 26 GuidedReview - Review AI-generated code before you sign your name to ithttps://github.com/nshntarora/guidedreview ↩ 回到正文 · back to text
  27. 27 DashClaw - policy and approval layer for unattended coding agentshttps://github.com/ucsandman/DashClaw ↩ 回到正文 · back to text
  28. 28 A local visual tool cli/mcp for agents to propose architecture changeshttps://github.com/luiscleto/WorkBraid ↩ 回到正文 · back to text
  29. 29 Awareness-Market - Open-source agent memory layer, 96% on LongMemEval, local-firsthttps://github.com/everest-an/Awareness-Market ↩ 回到正文 · back to text
  30. 30 MathKernel: An evidence-aware multi-engine mathematics kernel and MCP serverhttps://github.com/Staatsgeheim/MathKernel ↩ 回到正文 · back to text
  31. 31 A Tesla ran a stop sign and killed a man, Full Self-Driving/Autopilot was onhttps://electrek.co/2026/09/07/tesla-driver-assist-stop-sign-buena-vista/ ↩ 回到正文 · back to text
  32. 32 Tesla killing Solar Roof is leaving installers with six-figure losseshttps://electrek.co/2026/09/01/tesla-solar-roof-exit-installers-losses/ ↩ 回到正文 · back to text
  33. 33 This Month in Ladybird - August 2026https://ladybird.org/newsletter/2026-08-31/ ↩ 回到正文 · back to text
  34. 34 Scientists observe Einstein's gravity in the quantum worldhttps://www.ox.ac.uk/news/2026-08-28-scientists-observe-einsteins-gravity-in-the-quantum-world ↩ 回到正文 · back to text
  35. 35 Smartphone makers don't bother to comply with EU repairability requirementshttps://www.theregister.com/personal-tech/2026/09/07/smartphone-makers-dont-bother-to-comply-with-eu-repairability-requirements/5294532 ↩ 回到正文 · back to text
  36. 36 Switzerland's Federal Government Is Replacing Microsoft on 3k Computershttps://itsfoss.com/news/switzerland-replace-microssoft-pilot/ ↩ 回到正文 · back to text
  37. 37 Research acceleration: The view inside OpenAIhttps://openai.com/index/research-acceleration-view-inside-openai ↩ 回到正文 · back to text
  38. 38 PostgreSQL 19 Interactive Tourhttps://victoriametrics.com/blog/postgres-19/index.html ↩ 回到正文 · back to text
  39. 39 OpenAI brings back 5 hour limit for plus and business standard usershttps://news.ycombinator.com/item?id=49600233 ↩ 回到正文 · back to text
  40. 40 VMware migration reduces Tottenham Hotspur's licensing fees by 85 percenthttps://arstechnica.com/information-technology/2026/09/vmware-migration-reduces-tottenham-hotspurs-licensing-fees-by-85-percent/ ↩ 回到正文 · back to text
  41. 41 Initial effects of AI technology on employment look positivehttps://www.economist.com/finance-and-economics/2026/09/04/the-jobs-apocalypse-is-postponed-an-ai-jobs-boom-is-here ↩ 回到正文 · back to text
  42. 42 Nvidia's Jensen Huang says 'AGI has arrived' and congratulates OpenAIhttps://www.businessinsider.com/nvidia-jensen-huang-agi-openai-astra-ai-2026-9 ↩ 回到正文 · back to text
  43. 43 LG smart TVs caught logging audio with screen off and snooping on local deviceshttps://www.notebookcheck.net/LG-smart-TVs-caught-logging-audio-with-screen-off-and-snooping-on-local-devices.1391214.0.html ↩ 回到正文 · back to text
  44. 44 TiVo to charge money for skipping commercials in your own recordingshttps://cordcuttersnews.com/tivo-plans-to-end-free-automatic-commercial-skipping-in-november-tests-paid-premium-replacement-service/ ↩ 回到正文 · back to text
  45. 45 The Education of a Doomerhttps://borretti.me/article/the-education-of-a-doomer ↩ 回到正文 · back to text
  46. 46 Decapitating a MacBook (2025)https://mm-dev.rocks/series/decapitating-macbook-an-odyssey/ ↩ 回到正文 · back to text
  47. 47 Rebuilding a 1995 GPS Time Server so I don't get Telstra'dhttps://www.jeffgeerling.com/blog/2026/truetime-xl-gps-time-server-restomod/ ↩ 回到正文 · back to text
  48. 48 De-Brainrot Vacationshttps://devz.cl/posts/i-spent-my-vacations-de-brainrotting/ ↩ 回到正文 · back to text
  49. 49 Making a Python interpreter in 1024 byteshttps://austinhenley.com/blog/python1024.html ↩ 回到正文 · back to text
  50. 50 How well do agents use test/verification techniques?https://danluu.com/agentic-testing/ ↩ 回到正文 · back to text
  51. 51 Rust debugging survey 2026 resultshttps://blog.rust-lang.org/2026/09/07/rust-debugging-survey-2026-results/ ↩ 回到正文 · back to text
  52. 52 Programming is Arthttps://orchidfiles.com/programming-is-art/ ↩ 回到正文 · back to text
  53. 53 C Is Not a Low-Level Language (2018)https://queue.acm.org/doi/10.1145/3212477.3212479 ↩ 回到正文 · back to text
  54. 54 Bot Detection Without JavaScript: What My Blog Measuredhttps://gkoreli.com/how-i-separate-readers-from-bots-without-javascript ↩ 回到正文 · back to text
  55. 55 Demystifying complex configurationshttps://guix.gnu.org/blog/2026/demystifying-complex-configurations// ↩ 回到正文 · back to text
  56. 56 Data races and the limits of ThreadSanitizer in C and Gohttps://theconsensus.dev/p/2026/09/06/data-races-and-the-limits-of-threadsanitizer-in-c-and-go.html ↩ 回到正文 · back to text
  57. 57 Visualizing Rust's Vtables: How dyn Trait Works In Memoryhttps://sofiabelen.github.io/projects/visualizing-rusts-vtables-how-dyn-trait-works-in-memory/ ↩ 回到正文 · back to text
  58. 58 .gitignore Everything by Defaulthttps://packagemain.tech/p/gitignore-everything-by-default ↩ 回到正文 · back to text
  59. 59 There's No Limit to How Bad Code Can Gethttps://simonwillison.net/2026/Sep/6/theres-no-limit-to-how-bad-code-can-get/ ↩ 回到正文 · back to text
  60. 60 The-Swarm-Corporation/AutoHedgehttps://github.com/The-Swarm-Corporation/AutoHedge ↩ 回到正文 · back to text
  61. 61 sooryathejas/METATRONhttps://github.com/sooryathejas/METATRON ↩ 回到正文 · back to text
  62. 62 ahujasid/blender-mcphttps://github.com/ahujasid/blender-mcp ↩ 回到正文 · back to text
  63. 63 experientiallabs/experientialhttps://github.com/experientiallabs/experiential ↩ 回到正文 · back to text
  64. 64 aipoch/open-sciencehttps://github.com/aipoch/open-science ↩ 回到正文 · back to text
  65. 65 OpenWhispr/openwhisprhttps://github.com/OpenWhispr/openwhispr ↩ 回到正文 · back to text
  66. 66 nklmilojevic/sofkahttps://github.com/nklmilojevic/sofka ↩ 回到正文 · back to text
  67. 67 mixelpixx/Konnecthttps://github.com/mixelpixx/Konnect ↩ 回到正文 · back to text
  68. 68 pollen-robotics/microduck_rlhttps://github.com/pollen-robotics/microduck_rl ↩ 回到正文 · back to text
  69. 69 DietrichGebert/ponytailhttps://github.com/DietrichGebert/ponytail ↩ 回到正文 · back to text