📄 arXiv 论文速递

📅 2026-08-30cs.AI + cs.LG 最新提交 | DeepSeek 点评:这篇为什么重要

💡 该研究将智能体经验编译为持久化技能知识,解决技能自动发现与复用难题,推动智能体能力持续进化,对构建可扩展的自主系统具有重要意义。

Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities. Recent work automatically discovers such skills from agent exper

💡 针对软件工程任务中大规模轨迹数据构建成本高的问题,提出用更少轨迹实现更优性能的方法,提升大模型解决真实软件问题的效率,具有实际工程价值。

To improve large language models' ability to resolve real-world software issues, prior work has focused on constructing large-scale agent trajectory datasets and performing supervi

💡 现有代码审查基准多为静态,无法反映真实迭代过程。该工作提出动态基准,填补评估空白,推动代码审查模型向实用化发展,对软件质量保障至关重要。

In real-world software development, code review typically involves iterative interactions between developers and reviewers to improve software quality, making the process costly an

💡 针对LLM智能体在真实部署中的越狱风险,提出经验驱动的红队测试智能体,通过技能进化自动发现漏洞,增强产品级系统的安全性与鲁棒性。

LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks

Chemical reactions are fundamentally transformations in electron space, yet most machine learning approaches model them either through \textit{de novo} generation of product molecu

💡 在受治理组织中,LLM智能体需兼顾角色灵活演化与执行可审计性。该架构模式分离两者,解决安全与适应性冲突,为合规部署提供关键设计范式。

Large language model (LLM) agents in governed organizations must let the persona (instructions, tone, self-presentation) evolve freely, while keeping execution (stateful, audited w

Static scanners are increasingly used to identify executable or otherwise unsafe content in machine- learning artifacts, yet conventional evaluation metrics characterize only cases

Currently used sepsis severity indices rely on fixed variables and weights established decades ago, which are coarsely discretized and calibrated to a cohort that no longer reflect

Friend recommendation is inherently graph-structured: the relevance of a potential connection depends on multi-hop social context rather than user attributes alone. However, deploy

State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data

How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral

Clinical language models can achieve strong in-hospital accuracy yet fail under deployment shifts because they exploit note-specific artifacts (e.g., templates, separators, boilerp