arXiv: 2606.23449 (Technical Report); 源码分析基于 GitHub main 分支
AOHP 源码深度解读:Agent-Native OS 的架构实现与安全机制
AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction
涵盖 Agent 架构、形式化验证、具身智能与系统设计等领域的前沿论文解读。目前共 97 篇。
arXiv: 2606.23449 (Technical Report); 源码分析基于 GitHub main 分支
AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction
arXiv 2606.12344
Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks
arXiv 2605.22781
DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback
arXiv 2026
Trustworthy Software Project Generation: a Case Study with an Interactive Theorem Prover
arXiv
Qualixar OS: A Universal Operating System for AI Agent Orchestration
arXiv
Formalizing the Safety, Security, and Functional Properties of Agentic AI Systems
ICML 2025 (arXiv 2503.22738)
ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning
Codebase Analysis
DeerFlow Architecture Analysis: A Composable Agent Runtime
arXiv 2603.22928
SoK: The Attack Surface of Agentic AI -- Tools, and Autonomy
Codebase Analysis
Claw Agent Family Insights: A Comparative Study of Architecture, Security, and Extensibility
Preprints.org
Skills Are the New Apps – Now It's Time for Skill OS
arXiv
Agent Behavioral Contracts: Formal Specification and Runtime Enforcement for Reliable Autonomous AI Agents
Survey
Formal Verification Meets AI Agents: A Survey and Taxonomy
arXiv 2603.11088
The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey
ICLR 2026
PropagandaAI: An Analysis of Semantic Divergence in Large Language Models
arXiv (Position Paper)
Securing LLM Agents Need Intent-to-Execution Integrity
arXiv
A Framework for Formalizing LLM Agent Security
arXiv 2605.13044
No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills
ICML 2026 (PMLR 306)
Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction
arXiv
Uncovering Security Threats and Architecting Defenses in Autonomous Agents: A Case Study of OpenClaw
arXiv
Agent-Sentry: Bounding LLM Agents via Execution Provenance
arXiv:2510.11203
TraceAegis: Securing LLM-Based Agents via Hierarchical and Behavioral Anomaly Detection
arXiv:2604.04989
SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
arXiv 2604.01438
ClawSafety: "Safe" LLMs, Unsafe Agents
arXiv 2510.12985
SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents
ICSE 2026 (IEEE/ACM International Conference on Software Engineering)
AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents
arXiv
Temporal Logics for Hyperproperties
ICML 2019
Learning to Prove Theorems via Interacting with Proof Assistants
ACM Transactions on Software Engineering and Methodology (TOSEM)
Automatically Engineering Trusted Software: A Research Roadmap
Computer Science Review (Elsevier)
Linear Temporal Logic Symbolic Model Checking
arXiv
Probabilistic Model Checking: Applications and Trends
arXiv
FM-Agent: Scaling Formal Methods to Large Systems via LLM-Based Hoare-Style Reasoning
arXiv 2026
Intent Formalization: A Grand Challenge for Reliable Coding in the Age of AI Agents
CAV 2025
The Vampire Diary
CAV 2025
Lean-SMT: An SMT tactic for discharging proof goals in Lean
USENIX FAST '26
Sharpen the Spec, Cut the Code: A Case for Generative File System with SYSSPEC
J. ACM
ProofWright: Towards Agentic Formal Verification of CUDA
arXiv
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
IJRR 2025 / arXiv 2312.07843
Foundation Models in Robotics: Applications, Challenges, and the Future
arXiv 2507.10087
Foundation Model Driven Robotics: A Comprehensive Review
arXiv
An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
ICML 2026
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
arXiv
A Comprehensive Survey on World Models for Embodied AI
IEEE TNNLS / arXiv 2108.11544
Vision-Language Navigation: A Survey and Taxonomy
arXiv
Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
arXiv 2506.20966
Parallels Between VLA Model Post-Training and Human Motor Learning
IEEE TNNLS / arXiv 2405.14093
A Survey on Vision-Language-Action Models for Embodied AI
arXiv
Vision-Language-Action (VLA) Models: Concepts, Progress, Applications and Challenges
arXiv 2506.24044
A Survey on Vision-Language-Action Models for Autonomous Driving
arXiv 2507.01925 / PKU-PsiBot
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
arXiv 2502.06851(已撤回)
Survey on Vision-Language-Action Models
arXiv (Accepted to IJCAS)
SE(3)-Equivariant Robot Learning and Control: A Tutorial Survey
arXiv
Pure Vision-Language-Action (VLA) Models: A Comprehensive Survey
arXiv
A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI
arXiv
Neural Brain: A Neuroscience-inspired Framework for Embodied Agents
arXiv
Multimodal Perception for Goal-oriented Navigation: A Survey
arXiv
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
Science Robotics
A Review of Learning-based Dynamics Models for Robotic Manipulation
arXiv 2508.10399
Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
Blog Post
State of VLA Research at ICLR 2026
arXiv 2503.03464
Generative Artificial Intelligence in Robotic Manipulation: A Survey
arXiv 2408.11537
A Survey of Embodied Learning for Object-Centric Robotic Manipulation
arXiv 2502.15336
Exploring Embodied Multimodal Large Models: Development, Datasets, and Future Directions
arXiv
Embodied Intelligent Industrial Robotics: Concepts and Techniques
Annual Review / arXiv 2408.03539
Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes
arXiv
Diffusion Models for Robotic Manipulation: A Survey
arXiv
Dexterous Manipulation through Imitation Learning: A Survey
中国信息通信研究院 (CAICT)
具身智能发展报告 (2024 年) — 中国信息通信研究院
IEEE/ASME TMech / arXiv 2407.06886
Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
arXiv 2605.00080
World Model for Robot Learning: A Comprehensive Survey
NDSS
SAGA: A Security Architecture for Governing AI Agentic Systems
arXiv 2605.23904
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
arXiv
Self-Improvements in Modern Agentic Systems: A Survey
arXiv 2512.23738
Enforcing Temporal Constraints for LLM Agents
arXiv
Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety
IEEE S&P 2026
Towards Automating Data Access Permissions in AI Agents
arXiv 2607.01640
AgentFlow: Building Agent Dependency Graphs for Static Analysis of Agent Programs
arXiv (NVIDIA)
Polar: Agentic RL on Any Harness at Scale
arXiv 2607.01120
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents
arXiv preprint (2026)
HyPOLE: Hyperproperty-Guided Multi-Agent Reinforcement Learning under Partial Observation
arXiv preprint (Tencent HY LLM Frontier)
Harness Handbook: Making Evolving Agent Harnesses Readable, Navigable, and Editable
ICML 2026 (PMLR 306)
Automata-Conditioned Cooperative Multi-Agent Reinforcement Learning
arXiv 2606.25189
ActPlane: Programmable OS-Level Policy Enforcement for Agent Harnesses
ICLR 2025
Automated Design of Agentic Systems
arXiv
ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay
arXiv 2606.23525
Self-Compacting Language Model Agents
arXiv 2604.14228
Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
GitHub (colbymchenry/codegraph)
CodeGraph: A Semantic Code Knowledge Graph for AI Coding Assistants
arXiv 2505.24832
How much do language models memorize?
ICML 2016
Asynchronous Methods for Deep Reinforcement Learning
arXiv 2508.02736
AgentSight: System-Level Observability for AI Agents Using eBPF
GitHub Open Source
GenericAgent: A Minimalist Self-Evolving Agent Framework
GitHub Open Source
Pi: A Self-Extensible AI Coding Agent Framework
源码分析报告
Claw Agent Ecosystem: A Deep Architectural Comparison
Nous Research / GitHub
Hermes Agent - Self-Improving AI Agent Framework
ICLR 2026 (Oral)
GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
arXiv 2505.24832
How much do language models memorize?