主题 · Topics
主题 · Topics — 每日 Harness
按 Agent / Harness 主题跨刊期浏览全部内容——进入任一主题查看其全部条目并按日期筛选,或回到全部刊期按时间浏览。
共 7 个主题分组 · 21 个子主题 · 1533 条内容
1.1 推理与规划Reasoning & Planning 16 项
- XWM 2026-08-26github.com
- Active Inference as Context Acquisition for AI Agents 2026-08-22arxiv.org
- The Optimizer Is the Agent 2026-08-11arxiv.org
- Learning more about Claude’s mathematical capabilities 2026-08-11anthropic.com
- PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents 2026-07-24arxiv.org
1.2 测试时计算Test-time Compute 17 项
- Disagree to Explore, Agree to Converge: Adaptive MoE Routing for Test-Time Agentic Coding 2026-08-26arxiv.org
- Second Thought: Reasoning in Parallel as LLM Agents Act and Observe 2026-08-18arxiv.org
- Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things 2026-08-17simonwillison.net
- Test-Time Augmentation for LLMs 2026-08-12arxiv.org
- Test-Time Scaling in Reasoning LLMs 2026-08-06arxiv.org
2.1 Agent RL / 可验证奖励Agent RL / Verifiable Rewards 45 项
- radixark/miles 2026-09-06github.com
- MCP-Universe RL: A Framework for Training MCP Tool-Use Agents via Reinforcement Learning 2026-08-26arxiv.org
- Forgotten in Weights, Remembered by Agents: Trajectory-Aware Unlearning 2026-08-26arxiv.org
- SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning 2026-08-22arxiv.org
- How Much, Then Where 2026-08-11arxiv.org
2.2 蒸馏与压缩Distillation & Compression 13 项
- DeepSeek V4 Flash 56.8GB MoEspressoV2 2026-08-17huggingface.co
- Shoehorn 2026-08-15github.com
- Needle2 2026-08-12cactuscompute.com
- MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents 2026-08-11arxiv.org
- A Systematic Evaluation of Trajectory Data Curation for LoRA Fine-Tuning of Code Agents 2026-07-23arxiv.org
2.3 自进化Self-Evolution 21 项
- pi-autoresearch 2026-08-26github.com
- Auto-research with Codex: How I achieved a 232x Faster Kernel 2026-08-16sankalp.bearblog.dev
- PrimeIntellect-ai/prime-agent 2026-08-15github.com
- Mendel Gödel Machine 2026-08-12arxiv.org
- ADIAS: Automated Design of Interactive Agentic Systems 2026-08-11arxiv.org
2.4 合成数据与训练环境Synthetic Data & Environments 22 项
- Training Needs Trustworthy Worlds: Verified Environment Synthesis for Web Agents 2026-08-26arxiv.org
- MidTool: Mid-training Data Synthesis for Agentic Tool Use 2026-08-22arxiv.org
- EnvHarness: Awakening Static Worlds for Agent Learning 2026-08-22arxiv.org
- PhoneWorld: Scaling Phone-Use Agent Environments 2026-08-18arxiv.org
- newton-physics/newton 2026-08-16github.com
3.1 Agent 记忆Agent Memory 100 项
- EchoPath: Execution-Level Replayable Memory for GUI Agents 2026-09-18arxiv.org
- sillage 2026-08-26github.com
- FERAL-AI 2026-08-26github.com
- ECHO: Auditable Memory for Long-Horizon Agents 2026-08-26arxiv.org
- Ambient Context 2026-08-26github.com
3.2 上下文工程Context Engineering 99 项
- Portal by Spotify cut my Claude Code token usage by 90% 2026-09-06engineering.atspotify.com
- ST-Lite: Training-Free KV Cache Compression with Spatio-Trajectory Guidance for Long-Horizon GUI Agents 2026-08-28arxiv.org
- OpenWand 2026-08-28github.com
- From Inertia to Insight: Context Isolation for Robust Web Agents 2026-08-26arxiv.org
- Context as an Environment: Programmatic Context Management for Long-Horizon Agents 2026-08-26arxiv.org
3.3 检索与知识接地Retrieval / RAG 75 项
- claude-obsidian 2026-08-26github.com
- Read Less, Answer Better: A Training-Free Framework for Efficient Source-Grounded Agents 2026-08-26arxiv.org
- Oynix 2026-08-26oynix.dev
- Navigation-Informed Embeddings 2026-08-19arxiv.org
- OpenWebSearch 2026-08-18openwebsearch.ai
4.1 工具使用Tool Use 111 项
- Using Blender with coding agents on macOS 2026-09-06simonwillison.net
- Fast Cut Video tool for cutting video for Agents 2026-09-06github.com
- browser-use 2026-08-28github.com
- Telem 2026-08-28telem.ai
- BrowserSkill 2026-08-28github.com
4.2 技能系统Skills 74 项
- Cloudflare/Security-Audit-Skill 2026-09-18github.com
- humanlayer/skills 2026-09-06github.com
- earthtojake/text-to-cad 2026-09-06github.com
- scientific-agent-skills 2026-08-28github.com
- garden-skills 2026-08-28github.com
4.3 协议与互操作Protocols & Interop 62 项
- Yulin Local AWS Simulator 2026-09-06yulinsim.dev
- MCPay - spend authorization and per-call billing for MCP tools 2026-09-06github.com
- Concord 2026-08-28github.com
- Ratify Agent Relay Harness 2026-08-26github.com
- NeoBrowser 2026-08-19github.com
5.1 多智能体Multi-Agent 59 项
- tutti 2026-08-28github.com
- munder-difflin 2026-08-28github.com
- Diverse by Reasoning: Harnessing the Wisdom of LLM Crowds for Future Prediction 2026-08-28arxiv.org
- Architecture as a Capability Equalizer in Multi-Agent Systems 2026-08-26arxiv.org
- When Agents Coordinate 2026-08-19arxiv.org
5.2 工作流与控制流Workflows & Control 118 项
- Moadim.io - A scheduler for agents 2026-09-06moadim.io
- Valet 2026-08-28valet.dev
- AI Review Loops 2026-08-26kevinmahoney.co.uk
- Sollimann/bonsai 2026-08-19github.com
- Figranium 2026-08-19github.com
6.1 框架与脚手架Frameworks & Scaffolds 203 项
- MobileCode - OpenCode with Built-In iOS and Android Previews 2026-09-06github.com
- Ditch; Build multiple products at once 2026-09-06theditch.dev
- open-slide 2026-08-28github.com
- archify 2026-08-28github.com
- Z 2026-08-28github.com
6.2 执行环境与沙箱Execution & Sandboxing 72 项
- Lakr233/vphone-cli 2026-09-06github.com
- Ephemerals.dev - Disposable cloud dev environments for agents 2026-09-06ephemerals.dev
- AgentCloud 2026-08-28news.ycombinator.com
- MulmoTerminal 2026-08-26github.com
- Building an almost fully self-hosted, sandboxed, agentic software factory 2026-08-22blog.jakesaunders.dev
6.3 可观测性与调试Observability & Debugging 68 项
- AI handles incidents, engineers lose touch with their systems 2026-09-06sylvainkalache.com
- Unbox-AI 2026-08-26github.com
- apache/maka 2026-08-22github.com
- AgentSight 2026-08-22github.com
- When Agentic Executions Fail 2026-08-19arxiv.org
7.1 基准Benchmarks 90 项
- BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents 2026-09-18arxiv.org
- Can AI design circuit boards yet? 2026-09-06eebench.org
- SWE Refactor Bench: Whole-Repository Refactoring Under Hidden Tests 2026-08-26arxiv.org
- LitReview Arena: A Battle-Style Benchmark for Expert-Level Literature Reviews 2026-08-26arxiv.org
- K-Bench: Benchmarking Agentic Coding in Live Engineering Environments 2026-08-26arxiv.org
7.2 评测方法Eval Methodology 103 项
- Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds 2026-08-28arxiv.org
- AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval 2026-08-28arxiv.org
- There Is No Neutral Harness: Measurement Instability in LLM Evaluation 2026-08-26arxiv.org
- Agentic Scaffolding Can Induce Sycophancy in Language Models 2026-08-26arxiv.org
- Credit Without Ground Truth 2026-08-22arxiv.org
7.3 安全与攻防Security 145 项
- ForgeGuardian - Open-source software supply-chain security scanner 2026-09-06github.com
- The Hugging Face incident and the road ahead 2026-08-28openai.com
- Agentic Security: A Systems Framework for Tool-Using AI 2026-08-26arxiv.org
- AI-Infra-Guard 2026-08-26github.com
- Workspace Topology as an Attack Vector in Agentic Coding Assistants 2026-08-19arxiv.org
7.4 对齐与治理Alignment & Governance 20 项
- Fences, not sandboxes 2026-08-26yegge.ai
- Strengthening democratic oversight in national security 2026-08-19openai.com
- Pacing model development in an era of cyber-critical capabilities 2026-08-19openai.com
- Claude: System Prompts 2026-08-17platform.claude.com
- Hexis 2026-08-15github.com