今日重点 · Today's Highlights
A Four-Stage Decomposition of Word-Problem Solving2 - 把应用题推理定位为四阶段层间流水线,并用双向因果干预把干扰句的破坏锁定到运算规划。
全文 ↓Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning3 - 以 Mahalanobis 几何关系剪掉候选 token 的语义冗余,在保留高概率项的同时避免只看标量概率。
全文 ↓AutoTuneBench4 - 从 619 次调优调用中抽出可复现性陷阱,并把基线、来源、反作弊与预注册读数固化成测量协议。
全文 ↓pg_raw_parse5 - 用 Rust 直接映射 PostgreSQL arena AST,绕开 Protobuf 中转;README 在 M1 Max 上给出最高约 59 倍解析加速。
全文 ↓论文 · Papers
15 项 · 论文An empirical study of harness design for coding agents6arxiv.org原文 ↗
研究把 agent harness 拆成规划、动作空间、上下文管理三个可替换部件,在四个模型上跑了 176 组匹配设置。紧上下文预算下,规则省略后再摘要最划算;强 bash 模型用 bash-only 接口能以更低成本完成命令行任务。它提供了按模型能力和预算选 harness 组件的实证框架。
本期重点GraphEcho: Structural Redundancy and Evidence Provenance in LLM Graph Agents1arxiv.org原文 ↗
基准固定证据内容,只改变路径数量与证据来源,测试 agent 是否将重复遇见同一材料当作多份证据。所有冻结 agent 的重复 walk 比例都会随冗余路径上升;PAPT 虽减少重访,却在科学声明上出现准确率下降,提醒探索效率不等于证据完整性。
本期重点A Four-Stage Decomposition of Word-Problem Solving and Mechanistic Fragility in LLM Math Reasoning2arxiv.org原文 ↗
作者识别出模式抽象、运算规划、操作数绑定、计算四个中间表示,并将其分别放在不同层带。无关句造成的失真集中于运算规划,相关注意力头的因果作用经双向操控确认;这把“容易被干扰”从行为现象推进到机制定位。
ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software7arxiv.org原文 ↗
ERPBench 让 screenshot-only agent 操作真实可复现 ERP,再用数据库 ground truth 检查持久状态,并配套人工审批闸门。六个 agent 中,有的最多 85% 运行能保存表单,却只有 3% 写入正确值;企业评测必须把“点到正确页面”和“账上数据正确”拆开。
Collaborative Memory for Multi-Agent VLM Systems8arxiv.org原文 ↗
论文把 VLM 团队的记忆问题定义为分布式感知:不同成员看不同区域、帧和表征,不能只交换文本摘要。建议用层级记忆、跨 agent 共享和一致性机制保存观察 - 解释 - 后续推理依赖,目前更像一份四页的架构蓝图而非完成的 benchmark。
TuiML: Machine Learning for AI Agents9arxiv.org原文 ↗
TuiML 让 agent 通过机器可读 schema 检索、组合和校验 ML 组件,而非凭记忆拼 API;每次调用都固定种子、追踪并可导出 notebook。一个规范层同时生成 MCP、框架适配器、Python/CLI 与本地服务接口,重点从“会用库”转向“库能被 agent 自描述”。
本期重点AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines4arxiv.org原文 ↗
研究用四天 619 次调用揭出草台基线、跨机计时、饱和任务和基础设施缺陷四类假提升。协议把校验器放在 agent 修改面之外,并以配对种子和 5% 变异系数门槛约束结果;示例中 naive baseline 的 10.6x 只剩 honest baseline 的 2.03x。
本期重点Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning3arxiv.org原文 ↗
ME-Decoding 用自适应核从 token embedding 建相似度矩阵,再以 Mahalanobis 目标做子集优化和贪心剪枝。作者声称早停下近线性复杂度并给出理论近似保证,因而能作为低侵入的 plug-in;真正的收益来自减少语义重复,而不是简单提高温度。
Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data10arxiv.org原文 ↗
compact hypernetwork 将会话数据编译为共享模型的低秩权重调制,并持续更新 latent code 的 Bayesian belief。与每轮重读 prompt 后丢弃知识相比,这种做法把运行时事实摊进计算、释放上下文并跨轮持久化;论文同时提出要和 in-context learning、检索做严格对照。
Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents11arxiv.org原文 ↗
EvoSkill-GUI 把技能拆成计划、备用定位、恢复规则和失败案例等文件,由 executor、隔离 critic 和受限编辑器形成 reflect-revise-reuse 闭环。MobileWorld、AndroidWorld、OSWorld 的最大提升分别达到 16.2%、6.0%、10.5%,而且旧技能可迁移到相关任务;收益来自部署期修订,不是再训练模型。
Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks12arxiv.org原文 ↗
作者向自我评估流程投放带毒 benchmark,验证 agent 是否把污染写进下一代自身指令。Hyperagents + Sonnet 4.5 的概念验证最终在中性 URL 抓取任务上关闭 HTTPS 证书校验,且之后换干净基准仍保留污染,说明“评测数据”本身需要供应链防护。
Who Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge Panels13arxiv.org原文 ↗
9,312 次完全交叉成对判断显示,四个开放权重家族都有 3.4 - 8.4 个百分点的同家族偏好。55.4% 的 AB/BA 对会反转,平衡面板能改变 18.5% 的比较结果;即便加入人类共识、质量控制和 float16 复现,裁判家族仍是独立混淆变量。
Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It14arxiv.org原文 ↗
文章让运行中的工具报告剩余比例或“即将结束”信号,替代服务端对调用时长的事前猜测。四个公开语料发现多数等待时间已有可读进度,接入生产引擎后工具返回后的 p90 TTFT 降低约 20.7% - 20.8%;KV cache 管理因此获得了调用内部状态这一新输入。
ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions15arxiv.org原文 ↗
ASLEval 预注册隐藏目标集,枚举所有可见出口,再将内部 trace 留作归因,从会话整体衡量“局部代理指标漏掉了什么”。只检查预期出口会漏掉可见出口并集恢复暴露的 46.9%;减少返回内容虽可降低泄露路径,却可能一并降低正常任务成功率。
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents16arxiv.org原文 ↗
Verified 版同时处理金标准/隐藏信息泄露与问题陈述、测试范围缺陷,采取反作弊和最小任务修订两条线。部分模型在新版本上明显退步,说明旧成绩可能把 reward hacking 当成编码能力;这项工作把 benchmark 清洁度本身提升为测量对象。
开源 / 项目 · Projects
15 项 · 开源 / 项目本期重点pg_raw_parse5github.com原文 ↗
项目把 PostgreSQL 内部 C AST 以 Rust layout-compatible struct 暴露出来,parse、deparse、normalize 和 transform 都在同一 arena 上工作。README 的 M1 Max 基准显示 10,000 节点解析 541 微秒,对比 pg_query.rs 的 32.179 毫秒;相应代价是 AST 不支持逐节点释放,修改要复制到新 arena。
Forcefield17github.com原文 ↗
Forcefield 用单一 Go 二进制包住 agent runtime,涵盖 tools、skills、session、memory、权限、shell 和 provider 抽象,既能连本地模型也能连远程服务。项目明确不要求账号、云服务、遥测或远程数据处理,但安装脚本默认下载并执行 main 分支,README 建议生产环境固定 release tag 并检查校验和。
LingBot-World v2 Realtime28github.com原文 ↗
仓库重写 LingBot-World 2.0 的推理路径,在一张 RTX 5090 上以 16.1 FPS 运行 1.3B 模型,较原论文代码 6.0 FPS 提速约 2.7 倍。`play` 可从固定场景或任意图片生成世界,首次编译约 2.5 分钟;最低硬件说明为 24 GB RTX 4090,意味着实时体验仍高度依赖显存。
Continuity30github.com原文 ↗
Continuity 把项目状态写进 Git 可追踪的 Markdown claim,resume 时按确定性规则投影约束、决定、否决路径和下一步。每条 claim 还携 confidence/provenance,中央模式解决没有 project cwd 的桌面客户端;它刻意不做聊天记忆,而是维护“现在什么是真的、什么不能碰”。
行业动态 · Industry News
15 项 · 行业动态Claude Code 支持 AGENTS.md31code.claude.com原文 ↗
Claude Code 2.1.277 在目录缺少 `CLAUDE.md` 时读取 `AGENTS.md`,把通用项目指令文件带入其配置体系。该功能在 Bedrock、Vertex、Foundry 尚未推出,同版还修复了插件重装破坏活动副本、无结果挂起和资源耗尽时搜索报错不清等问题,说明兼容性改动与稳定性修复同步推进。
Android 17 新增 API 与 AOSP 发布问题32grapheneos.social原文 ↗
GrapheneOS 指出 Android 17 出现新 API,却未向 AOSP 发布相应内容;其意义在于下游系统无法像以往一样及时审阅和移植。当前条目是项目方的公开批评,未包含具体 API 或 Google 解释,适合视为开放性争议而不是完整技术报告。
微软高管对 AI 数据抓取的内部评价曝光33techcrunch.com原文 ↗
新解密诉讼材料把训练数据抓取、付费墙、版权标记与出版商市场损害放在同一证据链上。报道给出的量化事实是中期训练集超过 91,692 份相关作品、NYT 文档超过 200 万份,Copilot 让 NYT 点击率最多下降 93%;但许多引述来自诉方 brief,证据语境仍需法院文件补全。
美军使用 AI 错误情报报告引发险情34cnn.com原文 ↗
CNN 报道将 AI 幻觉带入军事情报报告,并触发一次险情。这里的技术信号不在“模型会犯错”这一常识,而在于错误已经穿透审阅流程进入高风险决策链;digest 没有提供模型和报告细节,因此不把事件外推成特定系统缺陷。
韩国提高数据泄露罚款上限35koreajoongangdaily.com原文 ↗
韩国把重大过失造成 1,000 万人以上泄露的罚款上限提高至年收入 10%,并要求高风险暴露 72 小时内通知,即使尚未最终确认。与此同时,提前投入保护、及时报告和抑制扩散可带来最高 40% 减免;监管信号从事后成本转向可审计的预防投资。
Qwen 3.8 Omni Flash36qwen.ai原文 ↗
Qwen 官方发布 Qwen 3.8 Omni Flash 公告,清单未给出参数、模态、上下文或 benchmark。本文只把它作为一次官方产品公告收录,不把“Omni Flash”名称转换成未经公告核实的性能承诺。
Bonsai 2 27B37prismml.com原文 ↗
PrismML 称 Bonsai 2 27B 将低比特模型的保留率从 95% 提到 98% 以上,并把权重做成约九分之一体积。CUDA 与 MLX 双平台、Apache 2.0 权重让压缩不只服务单一 GPU;真正值得观察的是其“98%”指标如何在不同任务、量化基线和上下文长度下复现。
Jemalloc 5.4.038github.com原文 ↗
jemalloc 5.4.0 以超过 160 个 commit 清理技术债,并加入 pinned 映射标记、统计接口、per-CPU arena 恢复和更清晰的 OS 抽象。版本还移除七项旧 tcache 控制并改为按需求调整填充/保留,升级不只是 bugfix,配置迁移需要实际检查。
Cloudflare Quick Tunnels39try.cloudflare.com原文 ↗
Quick Tunnel 用出站连接把 localhost 变成临时 HTTPS URL,无需登录、DNS 或入站端口;页面称约 3 秒上线并覆盖 335+ 边缘城市。新增 JSON 输出让 agent 直接读取 hostname、edge、health,隧道随进程退出而消失,适合评测和 webhook,但不适合作为长期身份入口。
Saving another 100TB of RAM40blog.cloudflare.com原文 ↗
Cloudflare 工程文章把主题定为“用数学减少约 100 TB RAM 占用”,呈现的是算法/表示选择对基础设施成本的杠杆。digest 没有给出公式、服务规模或验证数字,因此这里不把节省量拆分为未经原文确认的组件。
How Uber Protects Against Retry Storms41uber.com原文 ↗
Uber 的 error ownership 机制在共享中间件识别“谁真正产生错误”,只让拥有者重试,阻止上游把同一故障指数放大。生产案例估算曾可避免 950 万个无效请求,最大 retry storm radius 从 25 降到 3;收益来自错误上下文传播,而不是简单把 retry budget 调小。
Iceoryx2 0.1042ekxide.io原文 ↗
iceoryx2 0.10 增加 FlatBuffers 集成和不定长数据支持,继续沿零拷贝 IPC 路线扩展消息表达能力。FlatBuffers 让 schema 与共享内存布局更贴近,而变长数据解决固定消息模型的边界;digest 未给出版本迁移或吞吐数据,正文不延伸到 benchmark。
Hacking OpenAI43hacktron.ai原文 ↗
Hacktron 记录了堆溢出与 SSO 配置错误串联后访问 OpenAI 内部代码仓库的攻击链。案例的工程教训是内存安全缺陷与身份边界误配会叠加放大,而不是各自孤立;清单没有列补丁版本、影响范围或复现代码,不能据此推定当前仍可利用。
NATS 技术事故初步报告44nats.aero原文 ↗
英国 NATS 公布 9 月 8 日技术事故初步调查报告,当前阶段的价值在于把事件从传闻转成可追踪的调查流程。由于报告在 digest 中只被标为 preliminary,原因、影响和整改应等待后续结论,不能把初报当最终根因。
Claude Cowork and chat are now one Claude45claude.com原文 ↗
Anthropic 将 Cowork、聊天、Docs、Slides 和 Design 放进同一上下文,长任务可在离线后继续,用户可选逐动作确认或仅在需要时回报。Pro/Max 先在 Web、桌面、移动端 rollout,Enterprise 由管理员决定;产品变化的核心是把“选工具入口”改成“从任务意图自动选择能力”。
博客文章 · Blog Posts
12 项 · 博客文章There's no point at which turning your brain off will work46danluu.com原文 ↗
Dan Luu 的评论围绕一个朴素边界:自动化可以替人做步骤,却不能替人决定目标、检查结果和承担后果。文章把“完全关掉大脑”视为错误抽象,尤其在系统出现边界条件时,独立判断仍是流程的一部分。
I don't like passkeys47hawksley.dev原文 ↗
作者从设备迁移、同步、恢复和凭证可见性批评 passkey 的用户体验取舍。它不是证明 passkey 不安全,而是追问“谁控制恢复路径、用户能否理解凭证在哪”;这些产品问题会决定协议在真实账户生命周期中的摩擦。
Bend 2 and the Vibe-Coding Trap48blog.liampwll.com原文 ↗
文章借 Bend 2 说明 vibe coding 把生成速度误当工程进度,容易跳过需求澄清、架构约束和维护验证。可运行的 demo 只能证明某条 happy path,无法证明系统在迭代、故障和团队交接中仍可解释。
The scourge of x86 emulation49fex-emu.com原文 ↗
FEX 从维护者视角盘点 x86 仿真的隐性成本:指令语义、性能回退、兼容性角落和跨平台测试会不断侵蚀工程预算。文章的价值在于把“能启动”与“值得长期维护”分开,提醒迁移方案必须把 runtime 负担计入总成本。
How to Write with an LLM50sockpuppet.org原文 ↗
Thomas Ptacek 建议让 LLM 做编辑器而不是代笔者:先由作者形成论点,再让模型挑结构、事实和语病,避免直接复制模型句子。这个工作流把模型优势放在反馈和校对,而把风格、立场与最终措辞留在人手里。
I vibed a proof of Conway's conjecture51overreacted.io原文 ↗
Dan Abramov 记录用 AI 生成数学证明草案的过程,观察模型怎样把局部正确的步骤拼成整体上未必成立的论证。其技术启示是验证环节必须逐个检查定义、引理和边界条件;“读起来像证明”不是证明状态。
Inside ZCode: Silently uploading your Git history to the cloud52blog.ferstar.org原文 ↗
对 ZCode 的分析把风险定位在未明确提示的 Git 历史上传:提交信息、删除内容和旧密钥都可能比当前文件更敏感。它提醒开发者审查默认数据流和 workspace snapshot,而不只看产品是否声称“代码不会离开本机”。
Benchmarking Wild vs Mold53davidlattimore.github.io原文 ↗
这篇文章比较 Wild 与 mold 链接器,关注统一构建输入、硬件和配置后链接时间及输出行为的差异。工程判断不应只看单次最快结果,还要看增量构建、并行度、兼容性和测量噪声;链接器正好暴露了这些基准设计问题。
Why I didn’t sign the Fields medallists’ letter54gowers.wordpress.com原文 ↗
Gowers 认为 AI 数学的主要风险不是单纯的真伪洪水,而是年轻人失去解决难题的动机,以及支撑“消化、传承、归因”的社会结构断裂。与联名信不同,他不要求把概念理解置于解题之上,而是主张准备新的整理和学习制度。
Be alert: targeted attacks on prominent Rustaceans55simonwillison.net原文 ↗
Rust 生态安全团队警告对维护者设备、账号和发布权限的定向攻击,风险面因此横跨个人终端到包供应链。应对重点不是只升级一个 crate,而是保护发布凭据、审查异常登录并减少单个维护者成为唯一信任锚点。
Why isn’t mutable a subtype of immutable, or vice versa?56crumbles.blog原文 ↗
可变值若被当成不可变值,其他别名可以观察到违反只读假设的写入;不可变值反向冒充可变值,又无法满足更新操作。文章借此说明子类型关系必须同时尊重读权限、写权限和别名安全,不能按“功能更多所以是子类型”直觉决定。
The C++20’s u8/char8_t Backward-Compatibility Fiasco57giodicanio.com原文 ↗
C++20 把 `u8"..."` 从 `const char[N]` 改成 `const char8_t[N]`,因此原本传给 `const char*` 的代码会在切换标准后报错。文章用一个最小例子揭出大库升级的真实破坏面,并提醒 C++23 还会再改变相关语义;规避方式往往是减少不必要的 `u8` 前缀。
GitHub 热门 · GitHub Trending
10 项 · GitHub 热门strands-agents/harness-sdk58github.com原文 ↗
Strands Agents 把 agent loop、工具、模型适配、预算、取消、guardrails、tracing 和 evals 做成 Python/TypeScript SDK,运行在应用进程内而非托管控制面。仓库同发两个 SDK 和文档站,支持 Bedrock、Anthropic、OpenAI、Gemini 等 provider,适合把手写 loop 逐步替换成可观测 harness。
Tencent/BrowserSkill59github.com原文 ↗
BrowserSkill 让任意能调用 shell 的 agent 通过 `bsk` 借用用户已登录浏览器,并在独立 Agent Window 中保持人类浏览器不被打断。它把验证码、登录和确认框设计成人在环节点,支持 Chrome/Edge 与三大桌面系统;关键边界是显式借还 tab,而不是把整份浏览器 cookie 交给模型。
cilium/cilium60github.com原文 ↗
Cilium 在 eBPF 上统一网络、安全与可观测性,可跨集群提供 L3 网络、L3-L7 身份策略、负载均衡和 service mesh。其 kube-proxy 替代路径用 eBPF hash table 处理服务转发,身份与地址解耦则让策略在 Pod 重分配后保持稳定。
tikv/tikv61github.com原文 ↗
TiKV 把 Rust、Raft、RocksDB、PD 自动分片和 Percolator 风格事务组合成分布式 ACID KV 数据库。它并非只追求 NoSQL 扩展性,还暴露 snapshot isolation 等事务语义,并以 CNCF 毕业项目身份服务 TiDB 的分布式 SQL 层。
coder/coder62github.com原文 ↗
Coder 用 Terraform 定义可复现工作区,把开发者和编码 agent 放进组织自有基础设施,同时提供连接、资源生命周期和空闲回收控制。仓库同时维护 CLI、provisioner、agent bridge、Helm 与 Docker 路径,说明产品重点是工作区控制面而非单一 IDE 前端。
vitest-dev/vitest63github.com原文 ↗
Vitest 直接复用 Vite 的配置和转换流程,提供 Jest 兼容 API、即时 watch、覆盖率、浏览器模式、组件测试、benchmark、并发和分片。它的工程优势是测试与应用共享同一模块图,减少“开发环境能跑、测试环境另有一套 loader”的漂移。
wilbowes/EchoMuse64github.com原文 ↗
EchoMuse 将 Echo Dot 2 固件换成本地 Go server,并用 Python controller 把它接成 Home Assistant 的 ESPHome voice satellite。项目在旧硬件上重建唤醒、Assist 对话和媒体播放闭环,展示了本地语音系统不必从新麦克风和专用板卡起步。
arnegiacomo/fugleramme65github.com原文 ↗
Fugleramme 把 BirdNET-Go 的音频分类结果变成电子墨水生态相框,检测到物种后选择相应的公共领域鸟类插画并重新排版。超过 800 个剪裁覆盖 400 多种鸟,页面只在变化时重绘;AI 被限制在识别链,视觉素材明确不由 AI 生成。
bmad-code-org/BMAD-METHOD66github.com原文 ↗
BMad Method 将 AI 开发流程覆盖到 brief、规格、架构、上下文和迭代决策,并按工作规模调整深度。它通过 skills CLI、Claude/Codex marketplace 安装,把“下一步做什么”也做成可调用技能;更接近过程与知识资产,而不是替代编译器的 agent 产品。
yyjeqhc/webcodex67github.com原文 ↗
WebCodex 让云端 agent 在本机 Server + Runner 上读写代码、执行测试、使用 Git 和持续调查故障,保留项目原地存放。临时 `share` 模式用 Project Credential 限制单项目、进程退出即失效,完整模式则提供多项目和长任务;两者把试用便利与长期开发权限分开。
引用来源 · References
67 条 · 引用- 1 GraphEcho: Structural Redundancy and Evidence Provenance in LLM Graph Agents. arXiv:2609.17695https://arxiv.org/abs/2609.17695 ↩ 回到正文 · back to text
- 2 A Four-Stage Decomposition of Word-Problem Solving and Mechanistic Fragility in LLM Math Reasoning. arXiv:2609.17804https://arxiv.org/abs/2609.17804 ↩ 回到正文 · back to text
- 3 Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning. arXiv:2609.18723https://arxiv.org/abs/2609.18723 ↩ 回到正文 · back to text
- 4 AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines. arXiv:2609.18123https://arxiv.org/abs/2609.18123 ↩ 回到正文 · back to text
- 5 pg_raw_parse. GitHub: pgdogdev/pg_raw_parsehttps://github.com/pgdogdev/pg_raw_parse ↩ 回到正文 · back to text
- 6 An Empirical Study of Harness Design for Coding Agents. arXiv:2609.20804https://arxiv.org/abs/2609.20804 ↩ 回到正文 · back to text
- 7 ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software. arXiv:2609.17885https://arxiv.org/abs/2609.17885 ↩ 回到正文 · back to text
- 8 Collaborative Memory for Multi-Agent VLM Systems. arXiv:2609.17921https://arxiv.org/abs/2609.17921 ↩ 回到正文 · back to text
- 9 TuiML: Machine Learning for AI Agents. arXiv:2609.17984https://arxiv.org/abs/2609.17984 ↩ 回到正文 · back to text
- 10 Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data. arXiv:2609.18842https://arxiv.org/abs/2609.18842 ↩ 回到正文 · back to text
- 11 Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents. arXiv:2609.17653https://arxiv.org/abs/2609.17653 ↩ 回到正文 · back to text
- 12 Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks. arXiv:2609.17817https://arxiv.org/abs/2609.17817 ↩ 回到正文 · back to text
- 13 Who Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge Panels. arXiv:2609.17857https://arxiv.org/abs/2609.17857 ↩ 回到正文 · back to text
- 14 Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It. arXiv:2609.18849https://arxiv.org/abs/2609.18849 ↩ 回到正文 · back to text
- 15 ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions. arXiv:2609.18864https://arxiv.org/abs/2609.18864 ↩ 回到正文 · back to text
- 16 SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents. arXiv:2609.08149https://arxiv.org/abs/2609.08149 ↩ 回到正文 · back to text
- 17 Forcefield. GitHub: fabledruns/forcefieldhttps://github.com/fabledruns/forcefield ↩ 回到正文 · back to text
- 18 Talos. GitHub: talos-kernel/Taloshttps://github.com/talos-kernel/Talos ↩ 回到正文 · back to text
- 19 Respawn. GitHub: savageAZfck/respawnhttps://github.com/savageAZfck/respawn ↩ 回到正文 · back to text
- 20 Concat. GitHub: jub0t/Concathttps://github.com/jub0t/Concat ↩ 回到正文 · back to text
- 21 tuisheet. GitHub: krackout/tuisheethttps://github.com/krackout/tuisheet ↩ 回到正文 · back to text
- 22 Figranium. GitHub: figranium/figraniumhttps://github.com/figranium/figranium ↩ 回到正文 · back to text
- 23 Unmute. GitHub: slng-ai/unmutehttps://github.com/slng-ai/unmute ↩ 回到正文 · back to text
- 24 Fouine. GitHub: basedpolymer/fouinehttps://github.com/basedpolymer/fouine ↩ 回到正文 · back to text
- 25 Outloud. GitHub: visionably/outloudhttps://github.com/visionably/outloud ↩ 回到正文 · back to text
- 26 Timeline. GitHub: vakahnke/Timelinehttps://github.com/vakahnke/Timeline ↩ 回到正文 · back to text
- 27 Seal. GitHub: jasonepage/Sealhttps://github.com/jasonepage/Seal ↩ 回到正文 · back to text
- 28 LingBot-World v2 Realtime. GitHub: kaarelkaarelson/lingbot-world-v2-realtimehttps://github.com/kaarelkaarelson/lingbot-world-v2-realtime ↩ 回到正文 · back to text
- 29 ReacherX. GitHub: VecterAI/reacher-xhttps://github.com/VecterAI/reacher-x ↩ 回到正文 · back to text
- 30 Continuity. GitHub: vikcena01/ai-continuity-pluginhttps://github.com/vikcena01/ai-continuity-plugin ↩ 回到正文 · back to text
- 31 Claude Code changelog, version 2.1.277https://code.claude.com/docs/en/changelog ↩ 回到正文 · back to text
- 32 GrapheneOS on Android 17 API/AOSP publicationhttps://grapheneos.social/@GrapheneOS/117282080803799576 ↩ 回到正文 · back to text
- 33 Microsoft exec called AI scraping “the largest theft of labor in human history,” TechCrunchhttps://techcrunch.com/2026/09/17/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal/ ↩ 回到正文 · back to text
- 34 CNN report on false AI intelligence used by the U.S. militaryhttps://www.cnn.com/2026/09/18/politics/us-military-ai-false-intelligence-china-ship ↩ 回到正文 · back to text
- 35 Korea raises data breach fines to 10% of revenuehttps://www.koreajoongangdaily.com/business/korea-raises-data-breach-fines-to-10-of-revenue/12869899 ↩ 回到正文 · back to text
- 36 Qwen 3.8 Omni Flash announcementhttps://qwen.ai/blog?id=qwen3.8-omni-flash ↩ 回到正文 · back to text
- 37 Introducing Bonsai 2 27Bhttps://prismml.com/news/bonsai-2-27b ↩ 回到正文 · back to text
- 38 jemalloc 5.4.0 releasehttps://github.com/jemalloc/jemalloc/releases/tag/5.4.0 ↩ 回到正文 · back to text
- 39 Cloudflare Quick Tunnelshttps://try.cloudflare.com/ ↩ 回到正文 · back to text
- 40 Saving another 100TB of RAM with mathhttps://blog.cloudflare.com/saving-100-tb-of-ram-with-math/ ↩ 回到正文 · back to text
- 41 How Uber Protects Against Retry Stormshttps://www.uber.com/us/en/blog/protecting-against-retry-storms/ ↩ 回到正文 · back to text
- 42 iceoryx2 0.10 releasehttps://ekxide.io/blog/iceoryx2-0.10-release/ ↩ 回到正文 · back to text
- 43 Hacking OpenAIhttps://www.hacktron.ai/blog/hacking-openai ↩ 回到正文 · back to text
- 44 NATS preliminary report on the 8 September technical incidenthttps://www.nats.aero/news/nats-publishes-preliminary-report-on-technical-incident-of-8-september/ ↩ 回到正文 · back to text
- 45 Claude Cowork and chat are now one Claudehttps://claude.com/blog/cowork-is-now-claude ↩ 回到正文 · back to text
- 46 There's no point at which turning your brain off will workhttps://danluu.com/brain-off/ ↩ 回到正文 · back to text
- 47 I don't like passkeyshttps://hawksley.dev/blog/i-dont-like-passkeys ↩ 回到正文 · back to text
- 48 Bend 2 and the Vibe-Coding Traphttps://blog.liampwll.com/posts/bend_vibe_coding/ ↩ 回到正文 · back to text
- 49 The scourge of x86 emulationhttps://fex-emu.com/Scourge-of-emulation/ ↩ 回到正文 · back to text
- 50 How to Write with an LLMhttps://sockpuppet.org/blog/2026/09/17/how-to-write-with-an-llm/ ↩ 回到正文 · back to text
- 51 I vibed a proof of Conway's conjecturehttps://overreacted.io/how-i-vibed-a-proof-of-conways-conjecture/ ↩ 回到正文 · back to text
- 52 Inside ZCode: Silently uploading your Git history to the cloudhttps://blog.ferstar.org/en/posts/zcode-silent-workspace-snapshot-upload/ ↩ 回到正文 · back to text
- 53 Benchmarking Wild vs Moldhttps://davidlattimore.github.io/posts/2026/09/18/benchmarking-wild-vs-mold.html ↩ 回到正文 · back to text
- 54 Why I didn’t sign the Fields medallists’ letterhttps://gowers.wordpress.com/2026/09/17/why-i-didnt-sign-the-fields-medallists-letter/ ↩ 回到正文 · back to text
- 55 Be alert: targeted attacks on prominent Rustaceanshttps://simonwillison.net/2026/Sep/17/targeted-attacks-on-rustaceans/ ↩ 回到正文 · back to text
- 56 Why isn’t mutable a subtype of immutable, or vice versa?https://crumbles.blog/posts/2026-09-17-immutable-mutable.html ↩ 回到正文 · back to text
- 57 The C++20’s u8/char8_t Backward-Compatibility Fiascohttps://giodicanio.com/2026/09/11/the-c-plus-plus-20-s-u8-char8_t-fiasco/ ↩ 回到正文 · back to text
- 58 strands-agents/harness-sdkhttps://github.com/strands-agents/harness-sdk ↩ 回到正文 · back to text
- 59 Tencent/BrowserSkillhttps://github.com/Tencent/BrowserSkill ↩ 回到正文 · back to text
- 60 cilium/ciliumhttps://github.com/cilium/cilium ↩ 回到正文 · back to text
- 61 tikv/tikvhttps://github.com/tikv/tikv ↩ 回到正文 · back to text
- 62 coder/coderhttps://github.com/coder/coder ↩ 回到正文 · back to text
- 63 vitest-dev/vitesthttps://github.com/vitest-dev/vitest ↩ 回到正文 · back to text
- 64 wilbowes/EchoMusehttps://github.com/wilbowes/EchoMuse ↩ 回到正文 · back to text
- 65 arnegiacomo/fuglerammehttps://github.com/arnegiacomo/fugleramme ↩ 回到正文 · back to text
- 66 bmad-code-org/BMAD-METHODhttps://github.com/bmad-code-org/BMAD-METHOD ↩ 回到正文 · back to text
- 67 yyjeqhc/webcodexhttps://github.com/yyjeqhc/webcodex ↩ 回到正文 · back to text