Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation1 - 把 80 多个 agent benchmark 接入统一适配器,并从中审计出 82 个任务的可负担评测索引;8 个模型在 54 个 benchmark 上的最高配置通过率也只有 28.0%,为“评测规模”和“可运行性”同时设定了硬约束。
全文 ↓今日重点 · Today's Highlights
CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls2 - 将 provenance、授权和策略状态写成跨组件传递的可验证契约,在 2,560 个故障注入攻击实例中零有害外部效果,同时完成 700 个良性任务和升级 200 个含糊任务。
全文 ↓Agentic Context Cracking: Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data3 - 让文档阅读顺带生成面向未来查询的结构化数据;FanOutQA 在只增加一个相关问题时降本 53%,把一次性上下文开销转成可复用资产。
全文 ↓From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents4 - 用冻结快照和评估反馈持续版本化 GUI 操作技能;OSWorld 四个应用域在热身后提升 5.7 - 18.6 个百分点,但 GIMP 的跨任务检索和修订 churn 揭示了持久记忆的边界。
全文 ↓Keyclasp - Let agents use tokens without putting them in prompts5 - 通过本地加密 vault 和命令级注入让 agent 只接触 secret 名称;项目同时明确子进程仍可联网或写盘,适合作为泄漏面缩减层而非完整隔离方案。
全文 ↓论文 · Papers
15 项 · 论文A Removal Based Approach to Improve LLM Faithfulness at Test-Time6arxiv.org原文 ↗
论文把解释不忠实拆成遗漏真实影响的“不完整”和捏造影响的“不健全”,在推理时删除解释没有提及的输入概念后重新提问。两个数据集、多个模型家族和两种独立指标均显示其忠实度超过普通提示及“请保持忠实”提示;无需权重访问,代价是需要额外一次查询和概念归因步骤。
What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents7arxiv.org原文 ↗
研究固定 Aider、OpenHands、Qwen Code、SWE-agent 的任务记录,只比较 GRPO 在同一任务内按 harness 分组或跨 harness 混组。24,000 次密封 SWE-bench 评估里,执行 harness 把平均解决率从 2.14% 推到 9.27%,而分组规则在留出 harness 上仅 +0.25 个百分点且置信区间跨零,信用分配并非主要瓶颈。
MaxKernel: Agentic Kernel Generation for TPUs8arxiv.org原文 ↗
MaxKernel 用规划、实现、自调试、测试和 profiling 子代理串起三种 TPU kernel 搜索:人机协作、自动指标循环和图式全局探索。它在 50 项 JaxBench 任务及真实开源模型 workload 上达到专家手调实现的性能水平;论文的实际贡献在于把编译器反馈变成多代理共同的搜索信号,而不只是让模型一次性写 kernel。
SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents9arxiv.org原文 ↗
SiLR 把违规后的准入门建模为轨迹搜索,shadow-execute 每个提案,并按“有多少分支仍过载、各分支严重度”组成的乘积序保留结构信息。作者证明不存在对该序关系声音的标量替代;Gym-ANM 的多动作 episode 取得 21/21 恢复,对比终止式 0/21 和最佳标量门 9/21,说明阈值调参无法修复表示缺陷。
本期重点From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents4arxiv.org原文 ↗
这项工作把每轮 GUI 轨迹与评估器反馈写进可审计技能库,更新只在下一轮快照生效,因此能测量技能的增量而不是把能力提升归因于模型微调。四个 OSWorld 域在五轮空库热身后提升 5.7 - 18.6 个百分点;GIMP 中跨任务检索和 revision churn 也表明,重复接受编辑不等于找回最初任务。
Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing10arxiv.org原文 ↗
RSM-full 以余弦门控的 max-member merge 控制写入,以 atom-aware packer 组织读取,把“保留哪些记忆”和“如何装入窗口”作为一体化设计。AMA-Bench 的 4k 预算达到完整上下文 83% 质量、token 成本 32%;在约 2.6k - 5k 区间比 Online K-Means 高 3.5 - 6.0 个百分点(四种子,p<.001),是针对紧凑预算的实测 Pareto 点。
TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing11arxiv.org原文 ↗
TROVE 从已评估 workflow-search 轨迹提取复合技能和 outcome-conditioned 图,在线只编辑被新证据推翻的路线后缀,避免整条计划重算。代码、问答、数学三类任务都优于数据集级优化和图调度基线;消融把收益拆成复合技能的离线提升、插入动作的局部纠错,以及后缀替换带来的计算节省。
From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments12arxiv.org原文 ↗
综述按委托权限、时间持久性和环境耦合重新整理代理系统,并把模型、harness、环境三层分开,避免把接口扩张误写成自主性。基于截至 2026-08-31 的原始研究和官方规格,作者判断动作接口的证据最充分,而稳健完成、恢复、授权和独立验证仍没有同等强度的支持;这是一条对“能力叙事”降温的证据框架。
Substrate-Aware AI Agents: Execution Context as a First-Class Input13arxiv.org原文 ↗
研究把内存、运行时间、算力和操作约束显式写进规划状态,以高维两两距离代码任务检验不再“substrate blind”。给出 128 MB RAM、10 秒 wall-time 合约后,14 个对齐比较中 13 个峰值内存下降,三组模型平均耗时全部下降,最快达到 3.1 倍;约束信息改变的是实现选择,不只是最终筛选。
Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability14arxiv.org原文 ↗
研究将同一历史分别保存为原文、RAG、自然语言 NOTES 和固定 schema 知识图,测试写入模型更换后的读取、检索与修复。48 条随机答案码历史显示 KG-fixed 的准确率变化为 +0.0004±0.0020,而 NOTES 会随迁移方向出现 +9.91 或 −13.28 个百分点;记忆格式比“保留同一数据库”更决定升级风险。
CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents15arxiv.org原文 ↗
App-Forge、Task-Weave、Path-Steer 三段式管线把真实桌面应用转成共享状态的 GUI+CLI 环境,并把高效、验证过的混合路径用于后训练。9B 模型在 CUA-Verse 得分提升 39.3、步骤减少 37%、token 减少 60%;迁移到 OSWorld 仍有成功率 +16.8 个百分点和步骤 −57%,说明收益来自接口协同而非单纯视觉增强。
RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents16arxiv.org原文 ↗
RefactorPlatform 将模型、执行模式和提示具体度拆成可控轴,在隔离 workspace 中记录 diff、token、transcript,并用 AST 校验跨文件行为是否保持。100 个 RefactorBench 任务中,AST 感知分块比 token 窗口高 25 - 30%;精简检索单代理通过率 86% 胜过评测中的子代理 66%,检索的 token 开销被准确率收益抵消后,单次成功成本不变。
本期重点CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls2arxiv.org原文 ↗
CONTINUITY 用签名根授权、provenance commitment、角色绑定转移收据、类型化释放和 effect-bound permit,把每个组件的假设与保证连成端到端授权见证。参考实现覆盖 32 类故障;2,560 个参数化攻击实例无一提交有害外部效果,同时完成 700 个良性任务并升级 200 个含糊任务,重点在跨边界不丢失语义而非增加单点过滤器。
本期重点Agentic Context Cracking: Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data3arxiv.org原文 ↗
cracking 子代理从已加载文档上下文分叉,按观测到的问题决定抽取哪些结构,并对相关未来查询做投机性整理。论文指出在 FanOutQA 上理想结构化存储可便宜 28 倍;只扩展一个相关问题时降本 53% 且准确率保持,代价是要维护结构化结果的依据链和更新策略。
本期重点Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation1arxiv.org原文 ↗
适配器先通过代码审查和 parity experiment 把 80 多个 benchmark 统一进 Harbor,再用 54 个 benchmark、8 个模型分析能力与失败模式,最后过滤成 29 个 benchmark 的 82 个高质量任务。所有模型-harness 配置通过率低于 30%,最强 GPT-5.5+Codex 为 28.0%;这种刻意保持的难度为跨项目比较留下余量,也暴露出当前代理离稳健完成仍远。
开源 / 项目 · Projects
15 项 · 开源 / 项目本期重点Keyclasp - Let agents use tokens without putting them in prompts5github.com原文 ↗
Keyclasp 在本地加密 vault 保存凭据,命令执行前按名称注入实际 token,避免 secret 出现在提示词、代码文件和常规输出里。README 的安全边界很具体:被注入的子进程仍能把值发往网络、写入磁盘或交给同用户进程,因此它解决的是上下文暴露,不是沙箱问题。
Worldfixture - An OSS fake company your software can interact with17worldfixture.com原文 ↗
Worldfixture 用本地可运行的假公司模拟支付、邮件、存储和协作服务,让应用在无真实 SaaS 账号的情况下联调业务流程。项目把“外部依赖的替身”做成共享测试面,适合验证 agent 的多服务调用与错误处理,而不是只测单个 API mock。
Expose - a self-hosted tunneling tool like ngrok/localtunnel18github.com原文 ↗
Expose 是 Go 单二进制隧道,把 localhost 暴露为可分享 URL,主打 webhook 测试、移动调试和现场演示。零配置、免注册和自托管降低了把临时服务交给第三方的门槛;仓库同时保留 internal、文档和 GoReleaser,便于审查部署链。
MobileCode - OpenCode with React Native Previews19github.com原文 ↗
MobileCode 从 OpenCode 分叉,补上 iOS/Android 构建与 React Native 预览,让 agent 的编辑循环能看到真实移动端产物。它的价值不在新模型,而在把编译、运行和视觉回馈纳入 OpenCode 工作流,减少“代码看似完成但设备行为未知”的盲区。
OpenLivery - open-source multi-tenant WhatsApp agents for agencies20github.com原文 ↗
OpenLivery 提供代理机构管理客户 agent 的多租户 workspace、品牌门户、WhatsApp/网页聊天、知识库、自定义工具和人工转接。文档将 Docker 快速启动、租户隔离、密钥配置、备份升级分别列出,并允许自带 OpenAI/Anthropic key;它更像可运营的白标后台,而非单一聊天 demo。
Crew - Let Claude/Codex/OpenCode agents talk to each other21github.com原文 ↗
Crew 把三种 CLI 会话的状态、摘要和 transcript 尾部自动注入彼此上下文,并提供数秒内送达的 `crew send` agent-to-agent 邮件。全局 npm 安装会自动写入 Claude、Codex、OpenCode hook,支持 JSON 监控;它选择共享 checkout,因此并行修改的协调责任仍在用户和 agent。
Benzi - Code Intelligence Infrastructure for Frontier AI Models22github.com原文 ↗
Benzi 让 tree-sitter 编译器先解析每个文件、解析 import/继承/调用/引用,再从统一 query map 回答符号问题,而不是把仓库全文塞给模型。作者称十种语言共享同一图结构,并把调用流和数据流在调用点合并;这种 O(1) 查询定位的是跨文件事实,不能替代对行为语义的测试。
Coop - Isolated VM Environments for Running Claude Code and Codex23github.com原文 ↗
Coop 以 Rust CLI 管理 disposable VM,让 Claude Code/Codex 在获得 Docker、编译器和包管理器权限时与宿主机隔离。每个环境可重复创建和销毁,适合运行未知依赖、破坏性测试和无人值守任务;隔离粒度由 VM 配置决定,不能被描述成自动的应用级安全证明。
Engrim - A universal, local-first SQLite memory engine for AI CLIs24github.com原文 ↗
Engrim 把项目决策、约束和工作状态浓缩为约 4,000 字符的 episodic memory,客户端可经 MCP 自动写入,也能用 `engrim add` 手动保存。它把记忆从模型供应商解耦,支持 Antigravity、Claude Code、Cursor、Windsurf、Codex 之间切换;核心取舍是策展和检索质量,而非无限堆积聊天记录。
Chrome-bridge - let any AI agent drive your real logged-in Chrome25github.com原文 ↗
Chrome-bridge 通过一个扩展和零依赖 Node CLI 操作用户现有的登录 tab、SSO 和 profile,不要求 MCP-capable 客户端或重新启动浏览器。它提供 a11y snapshot、元素引用、CDP 截图和网络捕获,Chrome 117+/Node 18+ 即可运行;便利性同时意味着 agent 与真实账户权限处于同一浏览器边界。
GuidedReview - Review AI-generated code before you sign your name to it26github.com原文 ↗
GuidedReview 将本地 diff 和 GitHub PR 聚类成有顺序的 review units,LLM 只负责把 hunks 按 schema、逻辑、调用点、测试组织并加短评。CLI 与 Chrome 扩展直连用户模型和 GitHub,没有后端代审;代码来源始终是真实 diff,人的逐段确认仍是流程中的决策点。
DashClaw - policy and approval layer for unattended coding agents27github.com原文 ↗
DashClaw 把无人值守 agent 的远程审批、策略检查和执行证据集中在一层,并配套 SDK、MCP 配置、Dockerfile 与 smoke tests。它补足的是“动作能否执行、谁批准、留下什么证据”的治理链,而不是再造一个代码生成器;高风险策略的完整性仍取决于接入方是否覆盖所有执行入口。
A local visual tool cli/mcp for agents to propose architecture changes28github.com原文 ↗
WorkBraid 让人和 agent 在本地工作台中维护组件图、提交架构 proposal、比较版本并导出 PDF 报告,CLI 与 MCP 共用运行中的服务。源码运行要求 Go 1.26+、Node 24+,`workbraid mcp` 连接已有 UI;它把架构讨论从聊天文本变成可审阅的持久对象。
Awareness-Market - Open-source agent memory layer, 96% on LongMemEval, local-first29github.com原文 ↗
Awareness Local 以 Markdown 记忆卡为持久层,用 SQLite FTS5 与 embedding 混合检索,并通过 MCP 服务 Cursor、Claude Code、Copilot、Cline。一次 `npx @awareness.market/setup` 即可离线启动,还提供决策/风险知识卡和 Web dashboard;可解释的文本存储让人工纠正比黑盒向量库更直接。
MathKernel: An evidence-aware multi-engine mathematics kernel and MCP server30github.com原文 ↗
MathKernel 让 LLM 负责解析和解释,把精确、符号、形式化、区间证书和数值计算交给带 typed MathIR 的内核,并为每个结论记录 provenance 与 trust label。README 特别提醒“引擎一致不等于证明”,要求 claim-specific evidence bundle;MCP server、Python 库和便携产物把同一证据链暴露给不同工作流。
行业动态 · Industry News
14 项 · 行业动态A Tesla ran a stop sign and killed a man, Full Self-Driving/Autopilot was on31electrek.co原文 ↗
Electrek 报道一辆启用 FSD/Autopilot 的 Tesla 闯停牌并造成行人死亡;该条的技术问题集中在驾驶辅助开启状态如何进入事故记录与责任调查。新闻事实本身比“自动驾驶已成熟/失效”的泛化判断更窄,适合继续追踪官方调查材料。
Tesla killing Solar Roof is leaving installers with six-figure losses32electrek.co原文 ↗
Tesla 8 月停止供应 Solar Roof,转向传统光伏板,约七年在美国只完成 3,000 套、峰值每周 21 - 32 套,远低于每周 1,000 套承诺。安装商为多周培训、专用工具和施工团队投入数十万至数十万美元,如今还要面对未完工项目、零件和 25 年保修的不确定性;撤掉产品线的成本被转移给了合作网络。
This Month in Ladybird - August 202633ladybird.org原文 ↗
Ladybird 的 2026 年 8 月 newsletter 汇总浏览器项目月度开发进展,条目本身没有额外的技术数字。它的读法应是持续观察独立浏览器在引擎、平台和工程基础设施上的累积变化,而非把月报标题当成单次发布事件。
Scientists observe Einstein's gravity in the quantum world34ox.ac.uk原文 ↗
实验把超冷铷原子波分成一条被磁场托住、另一条自由下落的路径,再合并测量量子相位;测得相位与将等效原理应用于量子波的预测一致。牛津明确划出结论边界:这是首次直接测量自由落体的预测量子相位,不是证明引力本身量子化,也未达到检验宏观叠加崩溃所需质量和时长。
Smartphone makers don't bother to comply with EU repairability requirements35theregister.com原文 ↗
The Register 指出手机厂商对欧盟可维修性要求的执行并不充分,争议落在法规指标与实际拆修体验之间的落差。报道的意义不只在合规名单,也在于维修评分、零件供应和软件支持若不能联动,消费者看到的“可修”标签并不等于可持续使用年限。
Switzerland's Federal Government Is Replacing Microsoft on 3k Computers36itsfoss.com原文 ↗
瑞士联邦政府以 CHF 900 万启动 3,000 台工作站的 openDesk 试点,目标 2027 年底完成,期间与 Microsoft 365 并行;此前 172 人 PoC 对文档和邮件满意,但视频会议仍有技术限制。军方 Cyber Command 计划 2026 年 10 月前全量迁移,背后是外国访问风险、供应商依赖和授权成本,属于以渐进迁移换取数字主权的工程案例。
Research acceleration: The view inside OpenAI37openai.com原文 ↗
OpenAI 披露中位研究员到 8 月中每天消耗超过 600 美元 API 推理,90 分位超过 7,000 美元,组织总体达到每个人类工作日 3.1 个 agent-workday;但过去六个月成功的 4 - 8 小时任务中超过一半仍有人工介入。安全限制使 Astra GPU 配额再降 59.2%,其他模型配额增加 17.2% 抵消约 85%,数据同时显示自动化扩张与算力/安全治理瓶颈并存。
PostgreSQL 19 Interactive Tour38victoriametrics.com原文 ↗
VictoriaMetrics 用交互式示例导览 PostgreSQL 19 的新特性,重点是让读者在可执行上下文中理解行为变化,而非只读 release note。该条适合作为版本升级前的动手入口,具体 SQL 和执行器差异应结合官方文档与自己的 workload 验证。
OpenAI brings back 5 hour limit for plus and business standard users39news.ycombinator.com原文 ↗
Hacker News 讨论 OpenAI Plus 与 Business Standard 恢复五小时使用时长限制,帖子体现的是用户对配额政策变化的即时反馈,而非完整官方公告。对依赖长会话或并发 agent 的团队来说,限制的时间窗和套餐边界比“恢复限制”四个字更需要核对。
VMware migration reduces Tottenham Hotspur's licensing fees by 85 percent40arstechnica.com原文 ↗
热刺在三个月内从 VMware 迁到 HPE Morpheus VME/OpsRamp,授权费降幅超过 85%,但 CTO 强调把虚拟化、监控和 AI 运维放进统一界面才是主要收益。约 35 人的技术团队借自动发现和修复故障来放大人力;案例说明迁移回报同时由许可账单和运维编排能力决定。
Initial effects of AI technology on employment look positive41economist.com原文 ↗
The Economist 汇总的早期研究把 AI 对就业的初始影响描述为偏正面,标题所说的“jobs apocalypse postponed”强调观察窗口仍短。解读这类宏观结论需区分新增岗位、任务重分配和生产率收益,不能把短期净变化直接外推到长期职业结构。
Nvidia's Jensen Huang says 'AGI has arrived' and congratulates OpenAI42businessinsider.com原文 ↗
Jensen Huang 公开表示“AGI 已到来”并祝贺 OpenAI,这是一位产业领袖的判断性表态,不是带评测协议和可复现实验的科学结论。把它放入产业语境更合适:定义边界和证据标准仍未因一句庆祝性发言而统一。
LG smart TVs caught logging audio with screen off and snooping on local devices43notebookcheck.net原文 ↗
Notebookcheck 报道 LG 智能电视在屏幕关闭时记录音频并探测局域网设备,争点同时涉及麦克风生命周期、遥测披露和家庭网络最小权限。若调查成立,用户应把“待机”视为软件状态而非物理断电,并要求厂商提供可验证的开关与日志。
TiVo to charge money for skipping commercials in your own recordings44cordcuttersnews.com原文 ↗
TiVo 从 2026-11-02 起取消免费 SkipMode,改以 30 天试用后收费的 Premium Auto Commercial Skip 提供自动/一键跳过,价格尚未公布;旧录制的标记仍会工作,手动 30 秒跳过和快进不受影响。把曾经的核心 DVR 差异化改成订阅项,反映 Xperi 已从硬件制造转向电视系统和广告技术。
博客文章 · Blog Posts
15 项 · 博客文章The Education of a Doomer45borretti.me原文 ↗
这篇随笔回看作者的技术观念与悲观主义如何形成,重点在思想路径而非某个工程结论。它适合与“技术进步必然向好/必然失控”的单线叙事对照阅读,保留个人经验的局限也正是文本的一部分。
Decapitating a MacBook (2025)46mm-dev.rocks原文 ↗
作者为 Apple 平台开发买下屏幕损坏的 8GB M1 MacBook Air(约 £300),拆屏改成 headless 机器,绕开 macOS 虚拟机性能和云 CI 费用问题。文章逐段处理 USB-C 显示、供电、散热和远程访问,最终证明低功耗二手硬件也能支撑 Flutter 测试;价值在于记录真实摩擦而非展示一次性改装照片。
Rebuilding a 1995 GPS Time Server so I don't get Telstra'd47jeffgeerling.com原文 ↗
Geerling 用 Pi 5、GNSS HAT、gpsd 和 Chrony 把 1995 年 TrueTime XL-AK 恢复为 stratum-1 NTP 服务器,并复用原机 LCD/LED。固定风扇、force_turbo、隔离 PPS 中断和给晶振保温后,他在 6 - 12 小时测试中看到更稳定的频率,持续功耗约多 1W;这是一份把硬件历史转成可运行基础设施的实作笔记。
De-Brainrot Vacations48devz.cl原文 ↗
作者把假期设为减少数字媒体消费的自我实验,观察注意力和日常节奏如何变化。文章的技术性不在工具,而在把“少刷内容”当成可观察的行为干预;结论应理解为个人样本,不能直接当作普遍生产率方案。
Making a Python interpreter in 1024 bytes49austinhenley.com原文 ↗
文章挑战在 1024 字节内实现可运行的 Python 解释器,把语法、运行时和压缩策略同时推到极限。这个尺度迫使作者明确“语言实现最小闭环”需要哪些部件,也让字节级取舍比常规性能优化更可见;适合作为解释器结构的反向教学案例。
How well do agents use test/verification techniques?50danluu.com原文 ↗
Dan Luu 观察编码代理在真实流程中如何编写、运行和解释测试,问题焦点是验证行为是否跟得上生成速度。条目提醒读者区分“代理写出了测试”与“测试真的约束了结果”;评价 agent 时应看它是否主动复现失败、缩小假设并检查边界,而非只数测试文件。
Rust debugging survey 2026 results51blog.rust-lang.org原文 ↗
Rust 官方发布 2026 调试体验调查结果,面向工具链和开发反馈而非语言新特性。其价值在于把编译器错误、调试器和运行时诊断作为一个完整工作流来衡量,后续应结合原始问卷数据解读痛点排序。
Programming is Art52orchidfiles.com原文 ↗
文章讨论编程实践与艺术创作的相似处,关注选择、约束、表达和迭代如何共同形成作品。它并不提供工程指标,却能解释为什么代码质量常包含不可量化的结构感、节奏和取舍。
C Is Not a Low-Level Language (2018)53queue.acm.org原文 ↗
ACM Queue 文章重新追问 C 的“低级”标签,把抽象层次从语法表面移到编译器、内存模型和硬件行为的组合。这个视角提醒系统开发者:能写指针不等于直接控制机器,C 同样依赖大量未显式写出的运行时与平台约定。
Bot Detection Without JavaScript: What My Blog Measured54gkoreli.com原文 ↗
作者用 Cloudflare Worker/D1 记录 ASN、Fetch Metadata、Accept 头和 UA,强调“客户端是谁”与“是否有人阅读”不是同一问题。七天窗口里浏览器 UA 有 1,209 次页面事件、578 个日标识,Cloudflare 仪表盘只有 113 次页面、52 次访问;一个移动标识一秒抓 31 页,促使他改用统一时间窗、事件边界和可复核规则,而非寻找神奇比例。
Demystifying complex configurations55guix.gnu.org原文 ↗
Guix 博文从组织和调试角度解释复杂配置,目标是让声明式系统的依赖关系和失败位置更可见。它对应的工程问题不是“少写几行配置”,而是建立能被拆解、复现和局部验证的配置结构。
Data races and the limits of ThreadSanitizer in C and Go56theconsensus.dev原文 ↗
文章讨论 ThreadSanitizer 在 C 与 Go 数据竞争检测中的覆盖边界,提醒动态工具只能观察被执行到的路径和满足其插桩假设的内存访问。工程上仍需把静态分析、锁设计、模型检查和代码审查组合起来,不能把“TSan 没报错”当作无竞争证明。
Visualizing Rust's Vtables: How dyn Trait Works In Memory57sofiabelen.github.io原文 ↗
文章先用 C++ 虚函数和 CRTP 建立对照,再解释 Rust 泛型的静态分发、`dyn Trait` 的 vtable 指针,以及每个 (Type, Trait) 对应一张表。它还说明 object safety 为何禁止返回 `Self` 或带泛型参数的方法;把抽象规则落到内存布局,能避免把 Rust 当成“换语法的 C++”。
.gitignore Everything by Default58packagemain.tech原文 ↗
作者建议用 `*` 默认忽略,再用 `!` 显式放行源码、README、go.mod 等文件,以免把 node_modules、IDE 文件和环境变量带进仓库。示例同时承认这不是普适规范,并给出 `git check-ignore -v` 的排查方式;它把版本控制从“事后清理”改成“允许清单”决策。
There's No Limit to How Bad Code Can Get59simonwillison.net原文 ↗
Simon Willison 认为直接重写技术债务系统很少成功:旧系统仍是移动靶,新团队又无法完整理解行为,最后常留下两个生产系统。更稳妥的路径是先补自动化测试,再通过小步迁移和针对性重构交付价值;这与等待“大爆炸”替换形成鲜明的风险差异。
GitHub 热门 · GitHub Trending
10 项 · GitHub 热门The-Swarm-Corporation/AutoHedge60github.com原文 ↗
AutoHedge 把 Director、Quant、Risk Management、Execution 四类 agent 串成分析、论证、仓位和下单流水线,输出结构化 JSON 并保留审计日志。当前 Solana 才支持全自动交易,Coinbase 仍在开发;仓库要求钱包私钥和模型 key,因此“风险优先”描述不能替代真实资金隔离与策略验证。
sooryathejas/METATRON61github.com原文 ↗
METATRON 在 Parrot OS 上完全离线运行,调用 nmap、whois、whatweb、curl、dig、nikto 做侦察,再让 Ollama 的 metatron-qwen 分析漏洞、提出利用与修复建议。MariaDB 五表保存完整扫描历史,agent loop 可在分析中追加工具,适合研究本地模型如何编排安全工具,但不应把模型建议当作授权测试结论。
ahujasid/blender-mcp62github.com原文 ↗
Blender MCP 通过 addon 与 MCP server 把自然语言映射为建模、场景创建和对象操作,支持任意 MCP 客户端选择的 LLM。快速开始是安装 uv、配置 server、装 Blender 插件;第三方集成的定位很明确,实际几何正确性仍由 Blender 场景和用户检查负责。
experientiallabs/experiential63github.com原文 ↗
Experiential 提供一个 OpenAI-compatible gateway,把托管、BYOK 和本地模型统一路由,并按用户/agent/用例限制可用模型和花费。它还能从生产流量学习,生成针对质量、速度和成本的专属 router;本地安装向导会显示默认 $50 command budget,说明治理面从 API 兼容延伸到消费控制。
aipoch/open-science64github.com原文 ↗
AIPOCH Open Science 把科学 agent、Python/R notebook、数据连接器和 provenance 放在一个跨平台本地工作区,支持从文献综述到模拟、可视化和报告生成。v0.26.0 增加 Slurm 远程执行与可去重的参考文献库,并把 notebook 调用呈现为摘要卡;可复现性由执行记录和引用链共同构成,而不只是保存代码。
OpenWhispr/openwhispr65github.com原文 ↗
OpenWhispr 用全局热键把语音转成文本、翻译、AI 命令和会议笔记,本地可选 Whisper/Parakeet,云模式则 BYOK;项目承诺无数据收集、无 telemetry。它还做 Zoom/Teams/FaceTime 检测、说话人分离和语义笔记搜索,但 Intel Mac 因 ONNX Runtime 不再提供 x86_64 包而缺少实时 speaker fingerprint。
nklmilojevic/sofka66github.com原文 ↗
Sofka 以 kube-rs、ratatui 和全异步 Rust 管线重做 Kubernetes TUI,通用对象渲染让 CRD 首日可用,Flux/Argo CD、Helm inspector、批量动作和后台 port-forward 都走原生 API。它还把“为什么坏了”做成确定性的证据视图,并强制生产删除护栏;当前 macOS/Linux 二进制尚未签名公证,部署时要处理 Gatekeeper。
mixelpixx/Konnect67github.com原文 ↗
Konnect 是 KiCAD 10 的单 Rust 二进制 MCP 插件,提供 222 个工具、21 个按需工具集,覆盖原理图、PCB 布局布线、ERC/DRC、设计审查、器件搜索和制造导出。它用官方 IPC(protobuf over NNG)替代旧的 Node→Python→SWIG 链,缩短调用路径并接入 KiCAD undo/redo;项目仍标注 beta,真实板级设计需要人工复核。
pollen-robotics/microduck_rl68github.com原文 ↗
microduck_rl 为约 800g、25cm 双足机器人提供 MuJoCo Warp/mjlab 的 PPO 环境,策略以 50Hz 训练、导出 ONNX 后部署到实体机。仓库把 BAM 执行器物理、域随机化、齿隙模拟和奖励设计一并记录;4096 环境约 1 - 2 小时可得到可用步态,但训练需要 CUDA GPU,sim2real 仍依赖真实硬件验证。
DietrichGebert/ponytail69github.com原文 ↗
Ponytail 通过跨客户端规则让编码 agent 少装依赖、少写包装代码,优先使用原生 HTML 和仓库已有能力。真实 FastAPI+React 仓库的 12 个 feature、Haiku 4.5 四次重复中,平均 LOC −54%、token −22%、成本 −20%、时间 −27%,安全检查保持 100%;效果在 date picker 等过度设计陷阱最明显,代码已经精简的任务几乎不变。
引用来源 · References
69 条 · 引用- 1 Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation. arXiv:2609.04298https://arxiv.org/abs/2609.04298 ↩ 回到正文 · back to text
- 2 CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls. arXiv:2609.05269https://arxiv.org/abs/2609.05269 ↩ 回到正文 · back to text
- 3 Agentic Context Cracking: Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data. arXiv:2608.31082https://arxiv.org/abs/2608.31082 ↩ 回到正文 · back to text
- 4 From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents. arXiv:2609.04869https://arxiv.org/abs/2609.04869 ↩ 回到正文 · back to text
- 5 Keyclasp - Let agents use tokens without putting them in promptshttps://github.com/AndreaCatalucci/keyclasp ↩ 回到正文 · back to text
- 6 A Removal Based Approach to Improve LLM Faithfulness at Test-Time. arXiv:2609.04343https://arxiv.org/abs/2609.04343 ↩ 回到正文 · back to text
- 7 What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents. arXiv:2609.04518https://arxiv.org/abs/2609.04518 ↩ 回到正文 · back to text
- 8 MaxKernel: Agentic Kernel Generation for TPUs. arXiv:2609.04523https://arxiv.org/abs/2609.04523 ↩ 回到正文 · back to text
- 9 SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents. arXiv:2609.04629https://arxiv.org/abs/2609.04629 ↩ 回到正文 · back to text
- 10 Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing. arXiv:2609.04915https://arxiv.org/abs/2609.04915 ↩ 回到正文 · back to text
- 11 TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing. arXiv:2609.05019https://arxiv.org/abs/2609.05019 ↩ 回到正文 · back to text
- 12 From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments. arXiv:2609.04894https://arxiv.org/abs/2609.04894 ↩ 回到正文 · back to text
- 13 Substrate-Aware AI Agents: Execution Context as a First-Class Input. arXiv:2609.05232https://arxiv.org/abs/2609.05232 ↩ 回到正文 · back to text
- 14 Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability. arXiv:2609.05339https://arxiv.org/abs/2609.05339 ↩ 回到正文 · back to text
- 15 CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents. arXiv:2609.05374https://arxiv.org/abs/2609.05374 ↩ 回到正文 · back to text
- 16 RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents. arXiv:2609.04898https://arxiv.org/abs/2609.04898 ↩ 回到正文 · back to text
- 17 Worldfixture - An OSS fake company your software can interact withhttps://worldfixture.com/ ↩ 回到正文 · back to text
- 18 Expose - a self-hosted tunneling tool like ngrok/localtunnelhttps://github.com/kernelshard/expose ↩ 回到正文 · back to text
- 19 MobileCode - OpenCode with React Native Previewshttps://github.com/hsandhu/mobilecode ↩ 回到正文 · back to text
- 20 OpenLivery - open-source multi-tenant WhatsApp agents for agencieshttps://github.com/sarrazola/openlivery ↩ 回到正文 · back to text
- 21 Crew - Let Claude/Codex/OpenCode agents talk to each otherhttps://github.com/0xmmo/crew/ ↩ 回到正文 · back to text
- 22 Benzi - Code Intelligence Infrastructure for Frontier AI Modelshttps://github.com/oooscoos/Benzi ↩ 回到正文 · back to text
- 23 Coop - Isolated VM Environments for Running Claude Code and Codexhttps://github.com/trailofbits/coop ↩ 回到正文 · back to text
- 24 Engrim - A universal, local-first SQLite memory engine for AI CLIshttps://github.com/timgordontg/engrim ↩ 回到正文 · back to text
- 25 Chrome-bridge - let any AI agent drive your real logged-in Chromehttps://github.com/siropkin/chrome-bridge ↩ 回到正文 · back to text
- 26 GuidedReview - Review AI-generated code before you sign your name to ithttps://github.com/nshntarora/guidedreview ↩ 回到正文 · back to text
- 27 DashClaw - policy and approval layer for unattended coding agentshttps://github.com/ucsandman/DashClaw ↩ 回到正文 · back to text
- 28 A local visual tool cli/mcp for agents to propose architecture changeshttps://github.com/luiscleto/WorkBraid ↩ 回到正文 · back to text
- 29 Awareness-Market - Open-source agent memory layer, 96% on LongMemEval, local-firsthttps://github.com/everest-an/Awareness-Market ↩ 回到正文 · back to text
- 30 MathKernel: An evidence-aware multi-engine mathematics kernel and MCP serverhttps://github.com/Staatsgeheim/MathKernel ↩ 回到正文 · back to text
- 31 A Tesla ran a stop sign and killed a man, Full Self-Driving/Autopilot was onhttps://electrek.co/2026/09/07/tesla-driver-assist-stop-sign-buena-vista/ ↩ 回到正文 · back to text
- 32 Tesla killing Solar Roof is leaving installers with six-figure losseshttps://electrek.co/2026/09/01/tesla-solar-roof-exit-installers-losses/ ↩ 回到正文 · back to text
- 33 This Month in Ladybird - August 2026https://ladybird.org/newsletter/2026-08-31/ ↩ 回到正文 · back to text
- 34 Scientists observe Einstein's gravity in the quantum worldhttps://www.ox.ac.uk/news/2026-08-28-scientists-observe-einsteins-gravity-in-the-quantum-world ↩ 回到正文 · back to text
- 35 Smartphone makers don't bother to comply with EU repairability requirementshttps://www.theregister.com/personal-tech/2026/09/07/smartphone-makers-dont-bother-to-comply-with-eu-repairability-requirements/5294532 ↩ 回到正文 · back to text
- 36 Switzerland's Federal Government Is Replacing Microsoft on 3k Computershttps://itsfoss.com/news/switzerland-replace-microssoft-pilot/ ↩ 回到正文 · back to text
- 37 Research acceleration: The view inside OpenAIhttps://openai.com/index/research-acceleration-view-inside-openai ↩ 回到正文 · back to text
- 38 PostgreSQL 19 Interactive Tourhttps://victoriametrics.com/blog/postgres-19/index.html ↩ 回到正文 · back to text
- 39 OpenAI brings back 5 hour limit for plus and business standard usershttps://news.ycombinator.com/item?id=49600233 ↩ 回到正文 · back to text
- 40 VMware migration reduces Tottenham Hotspur's licensing fees by 85 percenthttps://arstechnica.com/information-technology/2026/09/vmware-migration-reduces-tottenham-hotspurs-licensing-fees-by-85-percent/ ↩ 回到正文 · back to text
- 41 Initial effects of AI technology on employment look positivehttps://www.economist.com/finance-and-economics/2026/09/04/the-jobs-apocalypse-is-postponed-an-ai-jobs-boom-is-here ↩ 回到正文 · back to text
- 42 Nvidia's Jensen Huang says 'AGI has arrived' and congratulates OpenAIhttps://www.businessinsider.com/nvidia-jensen-huang-agi-openai-astra-ai-2026-9 ↩ 回到正文 · back to text
- 43 LG smart TVs caught logging audio with screen off and snooping on local deviceshttps://www.notebookcheck.net/LG-smart-TVs-caught-logging-audio-with-screen-off-and-snooping-on-local-devices.1391214.0.html ↩ 回到正文 · back to text
- 44 TiVo to charge money for skipping commercials in your own recordingshttps://cordcuttersnews.com/tivo-plans-to-end-free-automatic-commercial-skipping-in-november-tests-paid-premium-replacement-service/ ↩ 回到正文 · back to text
- 45 The Education of a Doomerhttps://borretti.me/article/the-education-of-a-doomer ↩ 回到正文 · back to text
- 46 Decapitating a MacBook (2025)https://mm-dev.rocks/series/decapitating-macbook-an-odyssey/ ↩ 回到正文 · back to text
- 47 Rebuilding a 1995 GPS Time Server so I don't get Telstra'dhttps://www.jeffgeerling.com/blog/2026/truetime-xl-gps-time-server-restomod/ ↩ 回到正文 · back to text
- 48 De-Brainrot Vacationshttps://devz.cl/posts/i-spent-my-vacations-de-brainrotting/ ↩ 回到正文 · back to text
- 49 Making a Python interpreter in 1024 byteshttps://austinhenley.com/blog/python1024.html ↩ 回到正文 · back to text
- 50 How well do agents use test/verification techniques?https://danluu.com/agentic-testing/ ↩ 回到正文 · back to text
- 51 Rust debugging survey 2026 resultshttps://blog.rust-lang.org/2026/09/07/rust-debugging-survey-2026-results/ ↩ 回到正文 · back to text
- 52 Programming is Arthttps://orchidfiles.com/programming-is-art/ ↩ 回到正文 · back to text
- 53 C Is Not a Low-Level Language (2018)https://queue.acm.org/doi/10.1145/3212477.3212479 ↩ 回到正文 · back to text
- 54 Bot Detection Without JavaScript: What My Blog Measuredhttps://gkoreli.com/how-i-separate-readers-from-bots-without-javascript ↩ 回到正文 · back to text
- 55 Demystifying complex configurationshttps://guix.gnu.org/blog/2026/demystifying-complex-configurations// ↩ 回到正文 · back to text
- 56 Data races and the limits of ThreadSanitizer in C and Gohttps://theconsensus.dev/p/2026/09/06/data-races-and-the-limits-of-threadsanitizer-in-c-and-go.html ↩ 回到正文 · back to text
- 57 Visualizing Rust's Vtables: How dyn Trait Works In Memoryhttps://sofiabelen.github.io/projects/visualizing-rusts-vtables-how-dyn-trait-works-in-memory/ ↩ 回到正文 · back to text
- 58 .gitignore Everything by Defaulthttps://packagemain.tech/p/gitignore-everything-by-default ↩ 回到正文 · back to text
- 59 There's No Limit to How Bad Code Can Gethttps://simonwillison.net/2026/Sep/6/theres-no-limit-to-how-bad-code-can-get/ ↩ 回到正文 · back to text
- 60 The-Swarm-Corporation/AutoHedgehttps://github.com/The-Swarm-Corporation/AutoHedge ↩ 回到正文 · back to text
- 61 sooryathejas/METATRONhttps://github.com/sooryathejas/METATRON ↩ 回到正文 · back to text
- 62 ahujasid/blender-mcphttps://github.com/ahujasid/blender-mcp ↩ 回到正文 · back to text
- 63 experientiallabs/experientialhttps://github.com/experientiallabs/experiential ↩ 回到正文 · back to text
- 64 aipoch/open-sciencehttps://github.com/aipoch/open-science ↩ 回到正文 · back to text
- 65 OpenWhispr/openwhisprhttps://github.com/OpenWhispr/openwhispr ↩ 回到正文 · back to text
- 66 nklmilojevic/sofkahttps://github.com/nklmilojevic/sofka ↩ 回到正文 · back to text
- 67 mixelpixx/Konnecthttps://github.com/mixelpixx/Konnect ↩ 回到正文 · back to text
- 68 pollen-robotics/microduck_rlhttps://github.com/pollen-robotics/microduck_rl ↩ 回到正文 · back to text
- 69 DietrichGebert/ponytailhttps://github.com/DietrichGebert/ponytail ↩ 回到正文 · back to text