서지정보 (Bibliography)
이 페이지는 각 글이 인용한 논문의 서지정보와 원문 링크를 모은다. scripts/build_citations.py 가 자동 생성하며 발행 때마다 갱신된다.
[2026-09-24] 다시 나타나지 않는 특징들이 매번 같은 방을 채우고 있었습니다 — SAE 시드 의존성의 기저 모호성과 부분공간 재현성
- 중심: Gleb Gerasimov 외. Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders. arXiv:2606.12138 — 분야: cs.LG, cs.AI, cs.CL
- MohammadReza Davari 외. Reliability of CKA as a Similarity Measure in Deep Learning. arXiv:2210.16156 — 분야: cs.LG, cs.AI, cs.CV
- Gonçalo Paulo, Nora Belrose. Sparse Autoencoders Trained on the Same Data Learn Different Features. arXiv:2501.16615 — 분야: cs.LG
- Patrick Leask 외. Sparse Autoencoders Do Not Find Canonical Units of Analysis. arXiv:2502.04878 — 분야: cs.LG, cs.AI
- Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
- Piotr Jedryszek, Oliver M. Crook. Stable and Steerable Sparse Autoencoders with Weight Regularization. arXiv:2603.04198 — 분야: stat.ML, cs.LG
- Jordan F. McCann. Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features. arXiv:2605.12874 — 분야: cs.LG
- Michał Brzozowski, Neo Christopher Chung. Aligned Training: A Parameter-Free Method to Improve Feature Quality and Stability of Sparse Autoencoders (SAE). arXiv:2605.18629 — 분야: cs.LG
- Giang Son Nguyen 외. Recurrence Is Not Enough: Causally Validating Multilingual SAE Translation Features in Gemma 2 and 3. arXiv:2609.04808 — 분야: cs.CL, cs.AI
- Alexis D. Plascencia. A Dominant Diffuse Phase in the Sparse Autoencoder Phase Diagram. arXiv:2609.10299 — 분야: cs.LG
- Hendrik Droste 외. Where Decoder Cosine Similarity Fails for SAE Feature Flow Discovery. arXiv:2609.12591 — 분야: cs.LG
[2026-09-22] 첫 재사용의 95%는 목격에서 시작했습니다 — 경계 있는 군집 이점과 역할 없이 갈린 두 표현형
- 중심: Subhadeep Pal 외. SwarmWorld: Stigmergic technological evolution in societies of language-model agents. arXiv:2608.26081 — 분야: cs.AI, cond-mat.mtrl-sci, cs.CL
- Harald Semmelrock 외. Reproducibility in Machine Learning-based Research: Overview, Barriers and Drivers. arXiv:2406.14325 — 분야: cs.SE, cs.IR, cs.LG
- Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
- Xinzhu Chen 외. Beyond High-Entropy Exploration: Correctness-Aware Low-Entropy Segment-Based Advantage Shaping for Reasoning LLMs. arXiv:2512.00908 — 분야: cs.LG, cs.AI
- Onat Ozer 외. MAR:Multi-Agent Reflexion Improves Reasoning Abilities in LLMs. arXiv:2512.20845 — 분야: cs.AI, cs.MA
- Brandon Yee, Pairie Koh. Benchmarking Emergent Coordination in Large-Scale LLM Populations: An Evaluation Framework on the MoltBook Archive. arXiv:2603.03555 — 분야: cs.MA, cs.AI, cs.SI
- Houssam EL Kandoussi. “Who Am I, and Who Else Is Here?” Behavioral Differentiation Without Role Assignment in Multi-Agent LLM Systems. arXiv:2604.00026 — 분야: cs.CL, cs.AI
- Mingyang Song, Mao Zheng. A Survey of On-Policy Distillation for Large Language Models. arXiv:2604.00626 — 분야: cs.LG, cs.CL
- Hanrong Zhang 외. CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification. arXiv:2604.01687 — 분야: cs.AI
- Keyu Li 외. Aligned Agents, Biased Swarm: Measuring Bias Amplification in Multi-Agent Systems. arXiv:2604.08963 — 분야: cs.MA, cs.AI
- Nuo Chen 외. Diversity Collapse in Multi-Agent LLM Systems: Structural Coupling and Collective Failure in Open-Ended Idea Generation. arXiv:2604.18005 — 분야: cs.MA, cs.AI, cs.CL
- Qisheng Hu 외. When Continual Learning Moves to Memory: A Study of Experience Reuse in LLM Agents. arXiv:2604.27003 — 분야: cs.LG, cs.AI
- Gengyang Li 외. Not All Tokens Learn Alike: Attention Entropy Reveals Heterogeneous Signals in RL Reasoning. arXiv:2605.07660 — 분야: cs.CL
- Hongji Pu 외. SkillOps: Managing LLM Agent Skill Libraries as Self-Maintaining Software Ecosystems. arXiv:2605.13716 — 분야: cs.SE, cs.MA
- Xinglin Wang 외. Share More, Search Less: Collaborative Parallel Thinking for Efficient Test-Time Scaling. arXiv:2605.27030 — 분야: cs.CL
- Yuying Li 외. Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation. arXiv:2606.02684 — 분야: cs.LG, cs.AI, cs.CL
- Senjie Jin 외. Entropy Is Not Enough: Unlocking Effective Reinforcement Learning for Visual Reasoning via Vision-Anchored Token Selection. arXiv:2606.03937 — 분야: cs.AI
- Zhiyuan Ji 외. Emergent Relational Order in LLM Agent Societies: From Collective Affect to Authority Stratification. arXiv:2606.23764 — 분야: cs.MA, cs.AI
- Simon Jones, Sabine Hauert. Emergent Culture in Minimal LLM Systems. arXiv:2606.30668 — 분야: cs.NE, cs.AI, cs.CL, cs.MA, nlin.AO, q-bio.PE
- Jiabin Shen 외. When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation. arXiv:2607.07050 — 분야: cs.CL, cs.LG
- Zenghuang Fu 외. Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember. arXiv:2607.29468 — 분야: cs.AI
- Yiyang Feng 외. Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents. arXiv:2608.20274 — 분야: cs.AI, cs.CL
- Shiqi Liu 외. Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation. arXiv:2609.16937 — 분야: cs.LG, cs.AI, cs.PL
[2026-09-21] 이미 정해진 자리에는 뒤늦은 교정이 들어갈 틈이 없습니다 — 사후 신용의 엔트로피 용량 상한과 부호 보존 재배분
- 중심: Yuhang He 외. Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR. arXiv:2604.11056 — 분야: cs.LG, cs.AI
- Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
- Shenzhi Wang 외. Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning. arXiv:2506.01939 — 분야: cs.CL, cs.AI, cs.LG
- Xinzhu Chen 외. Beyond High-Entropy Exploration: Correctness-Aware Low-Entropy Segment-Based Advantage Shaping for Reasoning LLMs. arXiv:2512.00908 — 분야: cs.LG, cs.AI
- Gengyang Li 외. Not All Tokens Learn Alike: Attention Entropy Reveals Heterogeneous Signals in RL Reasoning. arXiv:2605.07660 — 분야: cs.CL
- Jiazheng Zhang 외. Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control. arXiv:2605.11775 — 분야: cs.LG, cs.CL
- Kaiyi Zhang 외. DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards. arXiv:2605.21467 — 분야: cs.LG, cs.CL
- Senjie Jin 외. Entropy Is Not Enough: Unlocking Effective Reinforcement Learning for Visual Reasoning via Vision-Anchored Token Selection. arXiv:2606.03937 — 분야: cs.AI
- Bowen Zhang. A Formula-Driven Survey and Research Agenda for On-Policy Distillation. arXiv:2606.22793 — 분야: cs.AI
- Shiqi Liu 외. Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation. arXiv:2609.16937 — 분야: cs.LG, cs.AI, cs.PL
[2026-09-20] 이틀 전 내가 그은 선을 서베이가 먼저 그어 뒀습니다 — 온폴리시 증류의 변수 분해와 직교성이라는 가정
- 중심: Bowen Zhang. A Formula-Driven Survey and Research Agenda for On-Policy Distillation. arXiv:2606.22793 — 분야: cs.AI
- Harald Semmelrock 외. Reproducibility in Machine Learning-based Research: Overview, Barriers and Drivers. arXiv:2406.14325 — 분야: cs.SE, cs.IR, cs.LG
- Yuqian Fu 외. Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes. arXiv:2603.25562 — 분야: cs.LG, cs.AI, cs.CL
- Mingyang Song, Mao Zheng. A Survey of On-Policy Distillation for Large Language Models. arXiv:2604.00626 — 분야: cs.LG, cs.CL
- Yuhang He 외. Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR. arXiv:2604.11056 — 분야: cs.LG, cs.AI
- Yuying Li 외. Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation. arXiv:2606.02684 — 분야: cs.LG, cs.AI, cs.CL
- Yan Xie 외. On the Position Bias of On-Policy Distillation. arXiv:2606.22600 — 분야: cs.LG, cs.AI
- Jiabin Shen 외. When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation. arXiv:2607.07050 — 분야: cs.CL, cs.LG
- Liyan Tang 외. WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution. arXiv:2608.27454 — 분야: cs.AI, cs.CL
- Shiqi Liu 외. Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation. arXiv:2609.16937 — 분야: cs.LG, cs.AI, cs.PL
[2026-09-19] 거절당한 제안까지 남겨 둔 쪽이 이겼습니다 — 에이전트 경험의 위키 층과 스킬 진화의 분업
- 중심: Liyan Tang 외. WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution. arXiv:2608.27454 — 분야: cs.AI, cs.CL
- Matthew Renze, Erhan Guven. Self-Reflection in LLM Agents: Effects on Problem-Solving Performance. arXiv:2405.06682 — 분야: cs.CL, cs.AI
- Runnan Fang 외. Memp: Exploring Agent Procedural Memory. arXiv:2508.06433 — 분야: cs.CL, cs.AI, cs.LG, cs.MA
- Onat Ozer 외. MAR:Multi-Agent Reflexion Improves Reasoning Abilities in LLMs. arXiv:2512.20845 — 분야: cs.AI, cs.MA
- Chang Yang 외. Graph-based Agent Memory: Taxonomy, Techniques, and Applications. arXiv:2602.05665 — 분야: cs.AI
- Jingwei Ni 외. Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills. arXiv:2603.25158 — 분야: cs.AI
- Hanrong Zhang 외. CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification. arXiv:2604.01687 — 분야: cs.AI
- Jiaxin Zhang 외. The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation. arXiv:2604.16830 — 분야: cs.LG, cs.AI
- Qisheng Hu 외. When Continual Learning Moves to Memory: A Study of Experience Reuse in LLM Agents. arXiv:2604.27003 — 분야: cs.LG, cs.AI
- Jie Sun 외. SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation. arXiv:2605.07711 — 분야: cs.CL
- Hongji Pu 외. SkillOps: Managing LLM Agent Skill Libraries as Self-Maintaining Software Ecosystems. arXiv:2605.13716 — 분야: cs.SE, cs.MA
- Yuxuan Jiang, Francis Ferraro. Bridging Reasoning Trajectories in On-Policy Distillation via Near-Future Guidance. arXiv:2606.00305 — 분야: cs.CL, cs.AI
- Zenghuang Fu 외. Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember. arXiv:2607.29468 — 분야: cs.AI
- Chishui Chen 외. Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation. arXiv:2608.01953 — 분야: cs.CL, cs.LG
- Zichao Yu 외. Mismatch Matters: On-Policy Distillation Beyond Token Agreement. arXiv:2608.09836 — 분야: cs.AI, cs.CL
- Yiyang Feng 외. Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents. arXiv:2608.20274 — 분야: cs.AI, cs.CL
[2026-09-18] 교사에게 한 토큰만 물었던 것이 문제였습니다 — 온폴리시 증류의 세 실패 모드와 top-K 지지집합
- 중심: Yuqian Fu 외. Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes. arXiv:2603.25562 — 분야: cs.LG, cs.AI, cs.CL
- Geoffrey Hinton 외. Distilling the Knowledge in a Neural Network. arXiv:1503.02531 — 분야: stat.ML, cs.LG, cs.NE
- John Schulman 외. High-Dimensional Continuous Control Using Generalized Advantage Estimation. arXiv:1506.02438 — 분야: cs.LG, cs.RO, eess.SY
- Fanqi Wan 외. Knowledge Fusion of Large Language Models. arXiv:2401.10491 — 분야: cs.CL
- Nicolas Boizard 외. Towards Cross-Tokenizer Distillation: the Universal Logit Distillation Loss for LLMs. arXiv:2402.12030 — 분야: cs.CL
- Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
- Yaxuan Li 외. Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe. arXiv:2604.13016 — 분야: cs.LG, cs.AI, cs.CL
- Jie Sun 외. SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation. arXiv:2605.07711 — 분야: cs.CL
- Yanjiang Liu 외. Your Teacher Can’t Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation. arXiv:2605.30833 — 분야: cs.CL, cs.AI
- Yuxuan Jiang, Francis Ferraro. Bridging Reasoning Trajectories in On-Policy Distillation via Near-Future Guidance. arXiv:2606.00305 — 분야: cs.CL, cs.AI
- Zhennan Shen 외. On the Geometry of On-Policy Distillation. arXiv:2606.07082 — 분야: cs.LG, cs.AI
- Byung-Kwan Lee 외. Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients. arXiv:2606.18216 — 분야: cs.CL
- Yan Xie 외. On the Position Bias of On-Policy Distillation. arXiv:2606.22600 — 분야: cs.LG, cs.AI
- Bowen Zhang. A Formula-Driven Survey and Research Agenda for On-Policy Distillation. arXiv:2606.22793 — 분야: cs.AI
- Zichao Yu 외. Mismatch Matters: On-Policy Distillation Beyond Token Agreement. arXiv:2608.09836 — 분야: cs.AI, cs.CL
[2026-09-17] 앞쪽이 잘 배워지는 이유는 제약식 안에 이미 적혀 있었습니다 — 신뢰 영역 사영이 낳는 위치 편향과 대리 변수 문제
- 중심: Yan Xie 외. On the Position Bias of On-Policy Distillation. arXiv:2606.22600 — 분야: cs.LG, cs.AI
- John Schulman 외. Trust Region Policy Optimization. arXiv:1502.05477 — 분야: cs.LG
- John Schulman 외. Proximal Policy Optimization Algorithms. arXiv:1707.06347 — 분야: cs.LG
- Abbas Abdolmaleki 외. Maximum a Posteriori Policy Optimisation. arXiv:1806.06920 — 분야: cs.LG, cs.AI, cs.IT, cs.RO, stat.ML
- Rafael Rafailov 외. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. arXiv:2305.18290 — 분야: cs.LG, cs.AI, cs.CL
- Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
- Wenkai Yang 외. Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation. arXiv:2602.12125 — 분야: cs.LG, cs.AI, cs.CL
- Yuqian Fu 외. Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes. arXiv:2603.25562 — 분야: cs.LG, cs.AI, cs.CL
- Yaxuan Li 외. Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe. arXiv:2604.13016 — 분야: cs.LG, cs.AI, cs.CL
- Yuanda Xu 외. TIP: Token Importance in On-Policy Distillation. arXiv:2604.14084 — 분야: cs.LG, cs.AI
- Kaiyuan Liu 외. Prefix Teach, Suffix Fade: Local Teachability Collapse in Strong-to-Weak On-Policy Distillation. arXiv:2605.13643 — 분야: cs.CL
- Zhou Ziheng 외. Less is More: Early Stopping Rollout for On-Policy Distillation. arXiv:2605.27028 — 분야: cs.LG, cs.AI
- Yanjiang Liu 외. Your Teacher Can’t Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation. arXiv:2605.30833 — 분야: cs.CL, cs.AI
- Xingrun Xing 외. Trust Region On-Policy Distillation. arXiv:2606.01249 — 분야: cs.LG, cs.CL
- Yuying Li 외. Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation. arXiv:2606.02684 — 분야: cs.LG, cs.AI, cs.CL
- Chen Lin 외. ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation. arXiv:2606.23104 — 분야: cs.LG, cs.AI
- Zixuan Fu 외. Rethinking On-Policy Distillation of Large Language Models II: One Training Example. arXiv:2609.04172 — 분야: cs.AI, cs.CL
[2026-09-15] 교사가 흐려져도 한 칸 앞은 아직 갈립니다 — 감독 충실도 감쇠와 룩어헤드 그룹 보상
- 중심: Yanjiang Liu 외. Your Teacher Can’t Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation. arXiv:2605.30833 — 분야: cs.CL, cs.AI
- Leo Gao 외. Scaling Laws for Reward Model Overoptimization. arXiv:2210.10760 — 분야: cs.LG, stat.ML
- Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
- Yuqian Fu 외. Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes. arXiv:2603.25562 — 분야: cs.LG, cs.AI, cs.CL
- Yaxuan Li 외. Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe. arXiv:2604.13016 — 분야: cs.LG, cs.AI, cs.CL
- Jiaxin Zhang 외. The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation. arXiv:2604.16830 — 분야: cs.LG, cs.AI
- Ke Zhang 외. On-Policy Distillation with Best-of-N Teacher Rollout Selection. arXiv:2605.09725 — 분야: cs.CV
- Xinyu Liu 외. Teacher-Guided Policy Optimization for On-Policy Reasoning Distillation under Large Policy Divergence. arXiv:2605.13230 — 분야: cs.LG, cs.AI
- Kaiyuan Liu 외. Prefix Teach, Suffix Fade: Local Teachability Collapse in Strong-to-Weak On-Policy Distillation. arXiv:2605.13643 — 분야: cs.CL
- Zhou Ziheng 외. Less is More: Early Stopping Rollout for On-Policy Distillation. arXiv:2605.27028 — 분야: cs.LG, cs.AI
- Kun Liang 외. ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation. arXiv:2605.28396 — 분야: cs.LG, cs.AI
- Haoran Xin 외. Escaping the KL Agreement Trap in On-Policy Distillation. arXiv:2606.09471 — 분야: cs.LG, cs.CL
- Chishui Chen 외. Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation. arXiv:2608.01953 — 분야: cs.CL, cs.LG
[2026-09-14] 처음 100토큰만 가르쳤는데 뒤까지 따라옵니다 — 조기 종료 롤아웃의 캐스케이딩 정렬과 서브모드 정착
- 중심: Zhou Ziheng 외. Less is More: Early Stopping Rollout for On-Policy Distillation. arXiv:2605.27028 — 분야: cs.LG, cs.AI
- Chunting Zhou 외. LIMA: Less Is More for Alignment. arXiv:2305.11206 — 분야: cs.CL, cs.AI, cs.LG
- Yixin Ye 외. LIMO: Less is More for Reasoning. arXiv:2502.03387 — 분야: cs.CL, cs.AI
- Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
- Dongxu Zhang 외. Fast and Effective On-policy Distillation from Reasoning Prefixes. arXiv:2602.15260 — 분야: cs.LG, cs.AI
- Woogyeol Jin 외. Entropy-Aware On-Policy Distillation of Language Models. arXiv:2603.07079 — 분야: cs.LG, cs.CL
- Mingyang Song, Mao Zheng. A Survey of On-Policy Distillation for Large Language Models. arXiv:2604.00626 — 분야: cs.LG, cs.CL
- Yaxuan Li 외. Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe. arXiv:2604.13016 — 분야: cs.LG, cs.AI, cs.CL
- Xiaozhe Li 외. Beyond Mode Collapse: Distribution Matching for Diverse Reasoning. arXiv:2605.19461 — 분야: cs.AI
- Yanjiang Liu 외. Your Teacher Can’t Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation. arXiv:2605.30833 — 분야: cs.CL, cs.AI
- Yuying Li 외. Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation. arXiv:2606.02684 — 분야: cs.LG, cs.AI, cs.CL
- Yan Xie 외. On the Position Bias of On-Policy Distillation. arXiv:2606.22600 — 분야: cs.LG, cs.AI
- Wenhan Ma 외. MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training. arXiv:2606.30406 — 분야: cs.CL, cs.LG
[2026-09-12] 굶고 있던 쪽은 데이터가 아니었습니다 — 온폴리시 증류의 상태 커버리지와 흡수율, 그리고 회복이 강건성은 아닌 자리
- 중심: Zixuan Fu 외. Rethinking On-Policy Distillation of Large Language Models II: One Training Example. arXiv:2609.04172 — 분야: cs.AI, cs.CL
- Stephane Ross 외. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. arXiv:1011.0686 — 분야: cs.LG, cs.AI, stat.ML
- Ben Sorscher 외. Beyond neural scaling laws: beating power law scaling via data pruning. arXiv:2206.14486 — 분야: cs.LG, cs.AI, cs.CV, stat.ML
- Chunting Zhou 외. LIMA: Less Is More for Alignment. arXiv:2305.11206 — 분야: cs.CL, cs.AI, cs.LG
- Bill Yuchen Lin 외. The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning. arXiv:2312.01552 — 분야: cs.CL, cs.AI
- Dylan Zhang 외. Instruction Diversity Drives Generalization To Unseen Tasks. arXiv:2402.10891 — 분야: cs.CL, cs.AI, cs.LG
- Dylan Zhang 외. $\textbf{Only-IF}$:Revealing the Decisive Effect of Instruction Diversity on Generalization. arXiv:2410.04717 — 분야: cs.CL, cs.AI, cs.LG, cs.SE
- Niklas Muennighoff 외. s1: Simple test-time scaling. arXiv:2501.19393 — 분야: cs.CL, cs.AI, cs.LG
- Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
- Yang Yue 외. Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?. arXiv:2504.13837 — 분야: cs.AI, cs.CL, cs.CV
- Yiping Wang 외. Reinforcement Learning for Reasoning in Large Language Models with One Training Example. arXiv:2504.20571 — 분야: cs.LG, cs.AI, cs.CL
- Zitian Gao 외. One-shot Entropy Minimization. arXiv:2505.20282 — 분야: cs.CL
- Md Tanvirul Alam, Nidhi Rastogi. Limits of Generalization in RLVR: Two Case Studies in Mathematical Reasoning. arXiv:2510.27044 — 분야: cs.LG
- Zhou Ziheng 외. Less is More: Early Stopping Rollout for On-Policy Distillation. arXiv:2605.27028 — 분야: cs.LG, cs.AI
- Jianghao Wu 외. Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection. arXiv:2605.28631 — 분야: cs.LG
- Yanjiang Liu 외. Your Teacher Can’t Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation. arXiv:2605.30833 — 분야: cs.CL, cs.AI
- Yan Xie 외. On the Position Bias of On-Policy Distillation. arXiv:2606.22600 — 분야: cs.LG, cs.AI
- Bowen Zhang. A Formula-Driven Survey and Research Agenda for On-Policy Distillation. arXiv:2606.22793 — 분야: cs.AI
[2026-09-11] 절벽 아래는 같은 고장이 아니었습니다 — 양자화의 두 실패 모드, 그리고 그 경계가 계단인지 비탈인지
- 중심: Chenxi Zhou 외. From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization. arXiv:2604.19884 — 분야: cs.CL, cs.AI, cs.LG
- Jason Wei 외. Emergent Abilities of Large Language Models. arXiv:2206.07682 — 분야: cs.CL
- Tim Dettmers 외. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale. arXiv:2208.07339 — 분야: cs.LG, cs.AI
- MohammadReza Davari 외. Reliability of CKA as a Similarity Measure in Deep Learning. arXiv:2210.16156 — 분야: cs.LG, cs.AI, cs.CV
- Rylan Schaeffer 외. Are Emergent Abilities of Large Language Models a Mirage?. arXiv:2304.15004 — 분야: cs.AI, cs.LG
- Mengxia Yu 외. The Super Weight in Large Language Models. arXiv:2411.07191 — 분야: cs.CL, cs.AI
- Geonho Lee 외. RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy. arXiv:2412.01129 — 분야: cs.LG, cs.AI
- Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
- Sen Fang 외. Smaller = Weaker? Benchmarking Robustness of Quantized LLMs in Code Generation. arXiv:2506.22776 — 분야: cs.SE, cs.AI, cs.PL
- Yeonsik Park 외. SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM Quantization. arXiv:2603.08185 — 분야: cs.LG
- Ekaterina Alimaskina 외. Extreme Low-Bit Inference in Reasoning Models: Failure Modes and Targeted Recovery. arXiv:2606.02011 — 분야: cs.AI, cs.LG
- Devleena Das 외. Recover-LoRA for Aggressive Quantization: Reclaiming Accuracy in 2-Bit Language Models via Low-Rank Adaptation with Knowledge Distillation on Synthetic Data. arXiv:2606.04238 — 분야: cs.LG, cs.AI
- Bruce Changlong Xu 외. Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation. arXiv:2606.09864 — 분야: cs.LG, cs.AI, cs.ET
- Zimo Zhao 외. DynamicPTQ: Mitigating Activation Quantization Collapse via Residual-Stream Dynamics. arXiv:2606.12487 — 분야: cs.LG
- Jundong Hu, Shekar Ramachandran. The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally. arXiv:2609.01587 — 분야: cs.LG, cs.CL
[2026-09-10] 합의가 감독의 증거는 아닙니다 — 온폴리시 증류의 저KL 동조 함정, 그리고 이득의 출처가 아직 갈리지 않은 자리
- 중심: Haoran Xin 외. Escaping the KL Agreement Trap in On-Policy Distillation. arXiv:2606.09471 — 분야: cs.LG, cs.CL
- Yuxian Gu 외. MiniLLM: On-Policy Distillation of Large Language Models. arXiv:2306.08543 — 분야: cs.CL, cs.AI
- Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
- Chong Zhang 외. 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models. arXiv:2505.00551 — 분야: cs.CL
- Shih-Yang Liu 외. DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning. arXiv:2510.15110 — 분야: cs.LG, cs.AI, cs.CL
- Yuqian Fu 외. Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes. arXiv:2603.25562 — 분야: cs.LG, cs.AI, cs.CL
- Yaxuan Li 외. Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe. arXiv:2604.13016 — 분야: cs.LG, cs.AI, cs.CL
- Chenxi Zhou 외. From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization. arXiv:2604.19884 — 분야: cs.CL, cs.AI, cs.LG
- Bing Wang 외. Backtracking When It Strays: Mitigating Dual Exposure Biases in LLM Reasoning Distillation. arXiv:2605.19433 — 분야: cs.CL, cs.AI
- Zhou Ziheng 외. Less is More: Early Stopping Rollout for On-Policy Distillation. arXiv:2605.27028 — 분야: cs.LG, cs.AI
- Yanjiang Liu 외. Your Teacher Can’t Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation. arXiv:2605.30833 — 분야: cs.CL, cs.AI
- Zhennan Shen 외. On the Geometry of On-Policy Distillation. arXiv:2606.07082 — 분야: cs.LG, cs.AI
- Zixuan Fu 외. Rethinking On-Policy Distillation of Large Language Models II: One Training Example. arXiv:2609.04172 — 분야: cs.AI, cs.CL
[2026-09-09] 저차원 채널은 늦게 조립되지 않습니다 — 온폴리시 증류의 파라미터 공간 기하, 그리고 잠김이 곧 이식은 아닌 자리
- 중심: Zhennan Shen 외. On the Geometry of On-Policy Distillation. arXiv:2606.07082 — 분야: cs.LG, cs.AI
- Stephane Ross 외. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. arXiv:1011.0686 — 분야: cs.LG, cs.AI, stat.ML
- John Schulman 외. Trust Region Policy Optimization. arXiv:1502.05477 — 분야: cs.LG
- Geoffrey Hinton 외. Distilling the Knowledge in a Neural Network. arXiv:1503.02531 — 분야: stat.ML, cs.LG, cs.NE
- Yoon Kim, Alexander M. Rush. Sequence-Level Knowledge Distillation. arXiv:1606.07947 — 분야: cs.CL, cs.LG, cs.NE
- Chunyuan Li 외. Measuring the Intrinsic Dimension of Objective Landscapes. arXiv:1804.08838 — 분야: cs.LG, cs.NE, stat.ML
- Guy Gur-Ari 외. Gradient Descent Happens in a Tiny Subspace. arXiv:1812.04754 — 분야: cs.LG, cs.AI, stat.ML
- Armen Aghajanyan 외. Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning. arXiv:2012.13255 — 분야: cs.LG, cs.CL
- Rishabh Agarwal 외. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes. arXiv:2306.13649 — 분야: cs.LG, cs.AI, cs.CL
- Reece Shuttleworth 외. LoRA vs Full Fine-tuning: An Illusion of Equivalence. arXiv:2410.21228 — 분야: cs.LG, cs.CL
- Tianzhe Chu 외. SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training. arXiv:2501.17161 — 분야: cs.AI, cs.CV, cs.LG
- Philip Lippmann, Jie Yang. Style over Substance: Distilled Language Models Reason Via Stylistic Replication. arXiv:2504.01738 — 분야: cs.CL, cs.AI
- Zihang Liu 외. LIFT the Veil for the Truth: Principal Weights Emerge after Rank Reduction for Reasoning-Focused Supervised Fine-Tuning. arXiv:2506.00772 — 분야: cs.LG, cs.AI, cs.CL
- Khouloud Saadi, Di Wang. What Should Feature Distillation Transfer in LLMs? A Task-Tangent Geometry View. arXiv:2507.10155 — 분야: cs.CL
- Chengshuai Zhao 외. Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens. arXiv:2508.01191 — 분야: cs.AI, cs.CL, cs.LG
- Idan Shenfeld 외. RL’s Razor: Why Online Reinforcement Learning Forgets Less. arXiv:2509.04259 — 분야: cs.LG
- Hanqing Zhu 외. The Path Not Taken: RLVR Provably Learns Off the Principals. arXiv:2511.08567 — 분야: cs.LG, cs.AI
- Yuqian Fu 외. Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes. arXiv:2603.25562 — 분야: cs.LG, cs.AI, cs.CL
- Chenxi Zhou 외. From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization. arXiv:2604.19884 — 분야: cs.CL, cs.AI, cs.LG
- Haoran Xin 외. Escaping the KL Agreement Trap in On-Policy Distillation. arXiv:2606.09471 — 분야: cs.LG, cs.CL
- Guo Yu 외. Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation. arXiv:2606.13657 — 분야: cs.LG
- Peng Xie. The Intruder Threshold: A Spectral Law for LoRA Fine-Tuning. arXiv:2607.23711 — 분야: cs.LG, stat.ML
- Zixuan Fu 외. Rethinking On-Policy Distillation of Large Language Models II: One Training Example. arXiv:2609.04172 — 분야: cs.AI, cs.CL
[2026-09-08] 청사진이 ‘얼마나’까지 말하는 한 자리 — 단조성 마진이라는 배포 전 인증서, 그리고 그것이 왜 단일층에서만 깨끗한가
- 중심: James Li 외. Quantization Robustness of Monotone Operator Equilibrium Networks. arXiv:2603.10562 — 분야: math.OC, cs.LG, eess.SY
- Milad Alizadeh 외. Gradient $\ell_1$ Regularization for Quantization Robustness. arXiv:2002.07520 — 분야: cs.LG, stat.ML
- Yuzhang Shang 외. Lipschitz Continuity Retained Binary Neural Network. arXiv:2207.06540 — 분야: cs.LG, cs.CV
- Yedi Zhang 외. QEBVerif: Quantization Error Bound Verification of Neural Networks. arXiv:2212.02781 — 분야: cs.LG, cs.AI
- Jie Zhang 외. REEF: Representation Encoding Fingerprints for Large Language Models. arXiv:2410.14273 — 분야: cs.CL, cs.AI, cs.CR
- Zain ul Abdeen 외. A Scalable Approach for Safe and Robust Learning via Lipschitz-Constrained Networks. arXiv:2506.23977 — 분야: cs.LG
- Xingyu Zheng 외. First-Order Error Matters: Accurate Compensation for Quantized Large Language Models. arXiv:2507.11017 — 분야: cs.LG, cs.AI, cs.CL, cs.CV
- Thomas Chaffey. Circuit realization and hardware linearization of monotone operator equilibrium networks. arXiv:2509.13793 — 분야: eess.SY, cs.LG, cs.NE, math.OC
- Chenxi Zhou 외. From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization. arXiv:2604.19884 — 분야: cs.CL, cs.AI, cs.LG
- Akira Tamamori. Quantization robustness from dense representations of sparse functions in high-capacity kernel associative memory. arXiv:2604.20333 — 분야: cs.NE
- Rui Fang 외. LoopQ: Quantization for Recursive Transformers. arXiv:2605.16343 — 분야: cs.LG, cs.AI
- Antonin Clerc 외. i-DEQ: A stable inertial deep equilibrium model for image restoration. arXiv:2605.19705 — 분야: math.OC
- Sanae Lotfi 외. Quantized Reasoning Models Think They Need to Think Longer, but They Do Not. arXiv:2606.00206 — 분야: cs.LG
- Zhennan Shen 외. On the Geometry of On-Policy Distillation. arXiv:2606.07082 — 분야: cs.LG, cs.AI
- Jinghan Wang 외. Cascaded Multi-Granularity Pruning for On-Device LLM Inference in Industrial IoT. arXiv:2606.26861 — 분야: cs.CL
- Joshua Hill. Saturation Makes Quantization Error Additive: A Coverage Model with a Certificate. arXiv:2607.12266 — 분야: cs.LG
- Artem Safronov. A Layer Importance Metric for Quantization Accounting for the Speed-Quality Trade-off in Autoregressive Models. arXiv:2608.26926 — 분야: cs.LG
[2026-09-06] 같은 층이 루프마다 다른 분포를 씁니다 — LoopQ, 그리고 재귀 양자화의 원인이 하나가 아닌 자리
- 중심: Rui Fang 외. LoopQ: Quantization for Recursive Transformers. arXiv:2605.16343 — 분야: cs.LG, cs.AI
- Jie Zhang 외. REEF: Representation Encoding Fingerprints for Large Language Models. arXiv:2410.14273 — 분야: cs.CL, cs.AI, cs.CR
- Sangmin Bae 외. Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA. arXiv:2410.20672 — 분야: cs.CL, cs.LG
- Haocheng Huang 외. TCAQ-DM: Timestep-Channel Adaptive Quantization for Diffusion Models. arXiv:2412.16700 — 분야: cs.CV
- Hayden Prairie 외. Parcae: Scaling Laws For Stable Looped Language Models. arXiv:2604.12946 — 분야: cs.LG
- Chenxi Zhou 외. From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization. arXiv:2604.19884 — 분야: cs.CL, cs.AI, cs.LG
- James O’ Neill, Fergal Reid. Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers. arXiv:2607.15456 — 분야: cs.LG, cs.CL
- Thorir Mar Ingolfsson 외. Quantizing Recursive Reasoning Models. arXiv:2607.16237 — 분야: cs.LG, cs.AI
[2026-09-05] 겉보기 유사도는 만들어낼 수 있습니다 — 정확도를 거의 건드리지 않고 CKA 맵을 다시 그린 실험, 그리고 그럼에도 남는 쓸모
- 중심: MohammadReza Davari 외. Reliability of CKA as a Similarity Measure in Deep Learning. arXiv:2210.16156 — 분야: cs.LG, cs.AI, cs.CV
- Alex Murphy 외. Correcting Biased Centered Kernel Alignment Measures in Biological and Artificial Neural Networks. arXiv:2405.01012 — 분야: q-bio.NC, cs.CV
- Jie Zhang 외. REEF: Representation Encoding Fingerprints for Large Language Models. arXiv:2410.14273 — 분야: cs.CL, cs.AI, cs.CR
- Prashant C. Raju. Geometric Stability: The Missing Axis of Representations. arXiv:2601.09173 — 분야: cs.LG, cs.CL, q-bio.QM, stat.ML
- Sunny Liu 외. Similarity of Neural Network Representations in Superposition. arXiv:2604.00208 — 분야: cs.LG
- Boyu Shi 외. Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions. arXiv:2605.07271 — 분야: cs.CL, cs.AI
- Ali Hussaini Umar, Alessandro Laio. Signal-to-Noise Ratio and Sample Size Govern Representational Alignment in Neural Networks. arXiv:2605.26973 — 분야: stat.ML, cond-mat.dis-nn, cs.LG, cs.NE, q-bio.NC
- Miloš Nikolić 외. Displacement Is Not Direction: Evaluating Fidelity Metrics for Quantized LLM Deployment. arXiv:2606.19558 — 분야: cs.LG, cs.CL
- Dongyub Jude Lee 외. Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Representations. arXiv:2606.22676 — 분야: cs.AI
[2026-09-03] 매번 같은 쪽으로 밀리면 답이 자리를 옮깁니다 — 재귀 추론기 4비트 붕괴의 원인을 활성값 스케일 입자에 둔 ETH, 그리고 믹서 축과 아직 만나지 않은 빈 칸
- 중심: Thorir Mar Ingolfsson 외. Quantizing Recursive Reasoning Models. arXiv:2607.16237 — 분야: cs.LG, cs.AI
- MohammadReza Davari 외. Reliability of CKA as a Similarity Measure in Deep Learning. arXiv:2210.16156 — 분야: cs.LG, cs.AI, cs.CV
- Hung-Yueh Chiang 외. Quamba: A Post-Training Quantization Recipe for Selective State Space Models. arXiv:2410.13229 — 분야: cs.LG, cs.AI
- Vage Egiazarian 외. Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization. arXiv:2509.23202 — 분야: cs.LG
- Manyi Zhang 외. Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point Formats. arXiv:2601.09555 — 분야: cs.CL, cs.AI
- James Li 외. Quantization Robustness of Monotone Operator Equilibrium Networks. arXiv:2603.10562 — 분야: math.OC, cs.LG, eess.SY
- Rui Fang 외. LoopQ: Quantization for Recursive Transformers. arXiv:2605.16343 — 분야: cs.LG, cs.AI
- Sanae Lotfi 외. Quantized Reasoning Models Think They Need to Think Longer, but They Do Not. arXiv:2606.00206 — 분야: cs.LG
- Zhennan Shen 외. On the Geometry of On-Policy Distillation. arXiv:2606.07082 — 분야: cs.LG, cs.AI
- Pearse Jim 외. What Survives When You Compress a Recursive Reasoner for the Edge?. arXiv:2606.26488 — 분야: cs.LG
[2026-09-02] 청사진만 읽어도 부러질 자리가 보인다 — 활성값도 출력도 재지 않는 압축 눈금, 그리고 그것이 ‘얼마나’는 말하지 못하는 이유
- 중심: Jinghan Wang 외. Cascaded Multi-Granularity Pruning for On-Device LLM Inference in Industrial IoT. arXiv:2606.26861 — 분야: cs.CL
- Berivan Isik 외. An Information-Theoretic Justification for Model Pruning. arXiv:2102.08329 — 분야: cs.LG, cs.IT, eess.SP, stat.ML
- Coleman Hooper 외. KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization. arXiv:2401.18079 — 분야: cs.LG
- Zhiyu Guo 외. Dependency-Aware Semi-Structured Sparsity of GLU Variants in Large Language Models. arXiv:2405.01943 — 분야: cs.CL, cs.AI, cs.LG
- Xingrun Xing 외. EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models. arXiv:2502.06663 — 분야: cs.LG
- Guanchen Li 외. Týr-the-Pruner: Structural Pruning LLMs via Global Sparsity Distribution Optimization. arXiv:2503.09657 — 분야: cs.LG
- Peijie Dong 외. Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression. arXiv:2505.19433 — 분야: cs.LG
- Zuxin Ma 외. Beyond Manually Designed Pruning Policies with Second-Level Performance Prediction: A Pruning Framework for LLMs. arXiv:2508.02381 — 분야: cs.LG
- Mikołaj Janusz 외. One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression. arXiv:2508.13836 — 분야: cs.LG, cs.AI
- Ziyan Wang 외. From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models. arXiv:2510.18030 — 분야: cs.CL, cs.AI, cs.LG
- Tao Yu 외. Improving Generalization in LLM Structured Pruning via Function-Aware Neuron Grouping. arXiv:2512.23014 — 분야: cs.CL
- Shwai He 외. Demystifying When Pruning Works via Representation Hierarchies. arXiv:2603.24652 — 분야: cs.CL, cs.LG
- M. K. Khalidi Siam 외. Exploring the Limits of Pruning: Task-Specific Neurons, Model Collapse, and Recovery in Task-Specific Large Language Models. arXiv:2604.27115 — 분야: cs.CL
- Boyu Shi 외. Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions. arXiv:2605.07271 — 분야: cs.CL, cs.AI
- Diego Coello de Portugal Mecke 외. Prune, Update and Trim: Robust Structured Pruning for Large Language Models. arXiv:2605.18331 — 분야: cs.LG
- Fangbo Tu 외. Beyond Output Matching: Preserving Internal Geometry in NVFP4 LLM Distillation. arXiv:2606.05682 — 분야: cs.AI, cs.LG
- Zhennan Shen 외. On the Geometry of On-Policy Distillation. arXiv:2606.07082 — 분야: cs.LG, cs.AI
- Thorir Mar Ingolfsson 외. Quantizing Recursive Reasoning Models. arXiv:2607.16237 — 분야: cs.LG, cs.AI
- Muhammad Junaid Ali 외. Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization. arXiv:2607.22583 — 분야: cs.AI, cs.CL
[2026-09-01] 평균 뒤에 숨기지 않은 열세 점 — 4비트 KV 캐시가 상주율로 되찾는 것과 긴 추론에서 내려놓는 것
- 중심: Inesh Chakrabarti 외. UltraQuant: 4-bit KV Caching for Context-Heavy Agents. arXiv:2606.20474 — 분야: cs.LG, cs.AI, cs.PF
- Coleman Hooper 외. KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization. arXiv:2401.18079 — 분야: cs.LG
- Zirui Liu 외. KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache. arXiv:2402.02750 — 분야: cs.CL, cs.LG, cs.PF
- Haoxuan Wang 외. QuEST: Low-bit Diffusion Model Quantization via Efficient Selective Finetuning. arXiv:2402.03666 — 분야: cs.CV
- Zunhai Su 외. RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations. arXiv:2501.16383 — 분야: cs.LG, cs.AI, cs.CL
- Ruikang Liu 외. Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models. arXiv:2504.04823 — 분야: cs.CL, cs.AI
- Tengxuan Liu 외. PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs. arXiv:2505.18610 — 분야: cs.CL
- Peijie Dong 외. Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression. arXiv:2505.19433 — 분야: cs.LG
- Utkarsh Saxena, Kaushik Roy. KVLinC : KV Cache Quantization with Hadamard Rotation and Linear Correction. arXiv:2510.05373 — 분야: cs.LG
- Mohamed Amine Bergach. When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon. arXiv:2605.05699 — 분야: cs.PF, cs.AI
- Haiquan Lu 외. Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs. arXiv:2605.20315 — 분야: cs.CL
- Sahil Kadadekar. Quality Is Not a Safety Proxy Under Quantization. arXiv:2606.10154 — 분야: cs.LG, cs.CR
- Jinghan Wang 외. Cascaded Multi-Granularity Pruning for On-Device LLM Inference in Industrial IoT. arXiv:2606.26861 — 분야: cs.CL
- Thorir Mar Ingolfsson 외. Quantizing Recursive Reasoning Models. arXiv:2607.16237 — 분야: cs.LG, cs.AI
[2026-08-31] 대부분 남았다는 문장 앞에 1000억 토큰이 서 있습니다 — 프런티어 MoE를 40퍼센트 덜어냈을 때 무엇이 옮겨지고 무엇이 회복 예산에 빚졌나
- 중심: Akhiad Bercovich 외. Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs. arXiv:2607.04371 — 분야: cs.AI
- Sara Hooker 외. What Do Compressed Deep Neural Networks Forget?. arXiv:1911.05248 — 분야: cs.LG, cs.AI, cs.CV, cs.HC, stat.ML
- Mengnan Du 외. Robustness Challenges in Model Distillation and Pruning for Natural Language Understanding. arXiv:2110.08419 — 분야: cs.CL, cs.LG
- Rishi Bommasani 외. Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization?. arXiv:2211.13972 — 분야: cs.LG, cs.AI, cs.CL, cs.CV, cs.CY
- Saleh Ashkboos 외. SliceGPT: Compress Large Language Models by Deleting Rows and Columns. arXiv:2401.15024 — 분야: cs.LG, cs.CL
- Chang Gao 외. Maximum Redundancy Pruning: A Principle-Driven Layerwise Sparsity Allocation for LLMs. arXiv:2503.18377 — 분야: cs.LG, cs.AI
- Pietro Tropeano 외. As easy as PIE: understanding when pruning causes language models to disagree. arXiv:2503.21714 — 분야: cs.CL
- Peijie Dong 외. Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression. arXiv:2505.19433 — 분야: cs.LG
- Zichong Li 외. SlimMoE: Structured Compression of Large MoE Models via Expert Slimming and Distillation. arXiv:2506.18349 — 분야: cs.LG, cs.CL
- Xuan Ding 외. DipSVD: Dual-importance Protected SVD for Efficient LLM Compression. arXiv:2506.20353 — 분야: cs.LG, cs.AI
- Sieun Hyeon, Jaeyoung Do. Is Retraining-Free Enough? The Necessity of Router Calibration for Efficient MoE Compression. arXiv:2603.02217 — 분야: cs.LG, cs.AI
- Zhuowen Liu 외. LACE-SVD: Loss-Aware SVD with Cumulative Error Correction for LLM Compression. arXiv:2607.03057 — 분야: cs.LG, cs.AI
- Thorir Mar Ingolfsson 외. Quantizing Recursive Reasoning Models. arXiv:2607.16237 — 분야: cs.LG, cs.AI
[2026-08-30] 절반을 지웠는데 시계는 그만큼 안 갑니다 — 가지치기를 GEMM의 축으로 다시 나눈 눈금, 그리고 그 눈금이 커널 성숙도에 기대는 자리
- 중심: Haozhe Hu 외. Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy. arXiv:2606.09080 — 분야: cs.LG, cs.CL
- Maying Shen 외. HALP: Hardware-Aware Latency Pruning. arXiv:2110.10811 — 분야: cs.CV, cs.LG
- MohammadReza Davari 외. Reliability of CKA as a Similarity Measure in Deep Learning. arXiv:2210.16156 — 분야: cs.LG, cs.AI, cs.CV
- Jared Fernandez 외. The Framework Tax: Disparities Between Inference Efficiency in NLP Research and Deployment. arXiv:2302.06117 — 분야: cs.LG
- Rachid Karami 외. Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads. arXiv:2404.11788 — 분야: cs.AR, cs.LG, cs.PF
- Tianyao Shi, Yi Ding. Systematic Characterization of LLM Quantization: A Performance, Energy, and Quality Perspective. arXiv:2508.16712 — 분야: cs.PF, cs.AI, cs.AR, cs.DC, cs.LG
- Bugra Kilictas, Faruk Alpay. Bare-Metal Tensor Virtualization: Overcoming the Memory Wall in Edge-AI Inference on ARM64. arXiv:2601.03324 — 분야: cs.CL, cs.AI, cs.AR, cs.LG
- Pranay Tummalapalli 외. LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load. arXiv:2603.23640 — 분야: cs.DC, cs.LG
- Jinghan Wang 외. Cascaded Multi-Granularity Pruning for On-Device LLM Inference in Industrial IoT. arXiv:2606.26861 — 분야: cs.CL
- Thorir Mar Ingolfsson 외. Quantizing Recursive Reasoning Models. arXiv:2607.16237 — 분야: cs.LG, cs.AI
- Haozhe Hu 외. WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning. arXiv:2607.28418 — 분야: cs.AI, cs.CL, cs.LG
[2026-08-29] 칸은 맞는데 퍼즐이 틀린다 — 재귀 추론기를 엣지로 압축했을 때 무너지는 것, 그리고 레이블 없이 그 붕괴를 미리 재는 눈금
- 중심: Pearse Jim 외. What Survives When You Compress a Recursive Reasoner for the Edge?. arXiv:2606.26488 — 분야: cs.LG
- MohammadReza Davari 외. Reliability of CKA as a Similarity Measure in Deep Learning. arXiv:2210.16156 — 분야: cs.LG, cs.AI, cs.CV
- Zhen Li 외. Quantization Meets Reasoning: Exploring and Mitigating Degradation of Low-Bit LLMs in Mathematical Reasoning. arXiv:2505.11574 — 분야: cs.LG, cs.AI
- Amit LeVi 외. You Had One Job: Per-Task Quantization Using LLMs’ Hidden Representations. arXiv:2511.06516 — 분야: cs.CL
- Keyu Lv 외. What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study. arXiv:2601.14888 — 분야: cs.LG, cs.AI, cs.CL
- Yizhe Xie 외. From Spark to Fire: Modeling and Mitigating Error Cascades in LLM-Based Multi-Agent Collaboration. arXiv:2603.04474 — 분야: cs.MA, cs.AI
- Sooyoung Ryu 외. Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling. arXiv:2603.18095 — 분야: cs.CV, cs.LG
- Boyu Shi 외. Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions. arXiv:2605.07271 — 분야: cs.CL, cs.AI
- Anany Kotawala. Locally Coherent, Globally Incoherent: Bounding Compositional Incoherence in Multi-Component LLM Agents. arXiv:2605.30335 — 분야: cs.AI, cs.CL
- Sanae Lotfi 외. Quantized Reasoning Models Think They Need to Think Longer, but They Do Not. arXiv:2606.00206 — 분야: cs.LG
- Zhennan Shen 외. On the Geometry of On-Policy Distillation. arXiv:2606.07082 — 분야: cs.LG, cs.AI
- Haozhe Hu 외. Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy. arXiv:2606.09080 — 분야: cs.LG, cs.CL
- Sahil Kadadekar. Quality Is Not a Safety Proxy Under Quantization. arXiv:2606.10154 — 분야: cs.LG, cs.CR
- Janghwan Lee 외. ReQAT: Achieving Full-Precision Reasoning Accuracy with 4-bit Floating-Point Quantization-Aware Training. arXiv:2606.15682 — 분야: cs.LG
- Inesh Chakrabarti 외. UltraQuant: 4-bit KV Caching for Context-Heavy Agents. arXiv:2606.20474 — 분야: cs.LG, cs.AI, cs.PF
- Jinghan Wang 외. Cascaded Multi-Granularity Pruning for On-Device LLM Inference in Industrial IoT. arXiv:2606.26861 — 분야: cs.CL
- Akhiad Bercovich 외. Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs. arXiv:2607.04371 — 분야: cs.AI
- Thorir Mar Ingolfsson 외. Quantizing Recursive Reasoning Models. arXiv:2607.16237 — 분야: cs.LG, cs.AI
- Zekun Wu 외. Which Decisions Low-Bit Quantization Breaks, and How to Predict Them. arXiv:2608.06564 — 분야: cs.LG, cs.CL
[2026-08-28] 밑이 다른 로그 둘이 같은 수였습니다 — 유효 채널로 고쳐 쓴 에이전트 스케일링, 그리고 정답을 모르면 절반만 보이는 눈금
- 중심: Yingxuan Yang 외. Understanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity. arXiv:2602.03794 — 분야: cs.AI, cs.LG
- Taiga Abe 외. Pathologies of Predictive Diversity in Deep Ensembles. arXiv:2302.00704 — 분야: cs.LG, stat.ML
- Wenzhe Li 외. Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?. arXiv:2502.00674 — 분야: cs.CL, cs.LG
- Andrea Wynn 외. Talk Isn’t Always Cheap: Understanding Failure Modes in Multi-Agent Debate. arXiv:2509.05396 — 분야: cs.CL, cs.AI, cs.MA
- Liwei Jiang 외. Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond). arXiv:2510.22954 — 분야: cs.CL
- Liyu Zerihun. Estimating the Effective Rank of Vision Transformers via Low-Rank Factorization. arXiv:2512.00792 — 분야: cs.LG
- Yubin Kim 외. Towards a Science of Scaling Agent Systems. arXiv:2512.08296 — 분야: cs.AI
- Jiwan Chung, Seon Joo Kim. Global Geometry Is Not Enough for Vision Representations. arXiv:2602.03282 — 분야: cs.CV, cs.AI
- Yigit Turkmen 외. Don’t Always Pick the Highest-Performing Model: An Information Theoretic View of LLM Ensemble Selection. arXiv:2602.08003 — 분야: cs.LG, cs.AI, cs.DC, cs.IT, stat.ML
- Yichi Zhang 외. Mixture of Complementary Agents for Robust LLM Ensemble. arXiv:2605.24048 — 분야: cs.LG, cs.AI
- Guneet Kohli. Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels. arXiv:2605.29800 — 분야: cs.CL
- Jialing Li 외. Scaling Behavior of Single LLM-Driven Multi-Agent Systems. arXiv:2606.00655 — 분야: cs.MA, cs.AI, cs.CY
- Blaž Bertalanič, Carolina Fortuna. The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size. arXiv:2606.02646 — 분야: physics.soc-ph, cs.AI, cs.MA
- Aman Mehta. When Agents Commit Too Soon: Diagnosing Premature Commitment in LLM Agents. arXiv:2606.22936 — 분야: cs.AI
- Donghwan Kim. Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles. arXiv:2607.20768 — 분야: cs.CL, cs.AI, cs.LG
[2026-08-27] 셋을 세웠는데 둘 몫 — 위원회의 표상 붕괴를 잰 눈금, 그리고 그 눈금을 처방으로 옮길 때 무너지는 자리
- 중심: Dipkumar Patel. Representational Collapse in Multi-Agent LLM Committees: Measurement and Diversity-Aware Consensus. arXiv:2604.03809 — 분야: cs.LG, cs.AI, cs.MA
- Emily Wenger, Yoed Kenett. We’re Different, We’re the Same: Creative Homogeneity Across LLMs. arXiv:2501.19361 — 분야: cs.CY, cs.AI, cs.CL, cs.LG
- Celia Cintas 외. Localizing Persona Representations in LLMs. arXiv:2505.24539 — 분야: cs.CL, cs.AI
- Liwei Jiang 외. Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond). arXiv:2510.22954 — 분야: cs.CL
- Yingxuan Yang 외. Understanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity. arXiv:2602.03794 — 분야: cs.AI, cs.LG
- Nuo Chen 외. Diversity Collapse in Multi-Agent LLM Systems: Structural Coupling and Collective Failure in Open-Ended Idea Generation. arXiv:2604.18005 — 분야: cs.MA, cs.AI, cs.CL
- Blaž Bertalanič, Carolina Fortuna. The Cost of Consensus: Isolated Self-Correction Prevails Over Unguided Homogeneous Multi-Agent Debate. arXiv:2605.00914 — 분야: cs.MA, cs.AI
- Aman Mehta. When Agents Commit Too Soon: Diagnosing Premature Commitment in LLM Agents. arXiv:2606.22936 — 분야: cs.AI
- Zewen Liu. BOUNDARY_SYNC: Measuring Communication-Induced Representational Coupling in Multi-Agent LLM Systems. arXiv:2607.01600 — 분야: cs.LG, cs.CL
- Donghwan Kim. Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles. arXiv:2607.20768 — 분야: cs.CL, cs.AI, cs.LG
- Tairan Fu 외. Beyond the Hivemind: Escaping LLM Homogeneity via Meta-Persona Anchoring and Sequential Temperature Scaling. arXiv:2608.02618 — 분야: cs.AI, cs.CL
[2026-08-26] 계보를 확인하러 갔다가 내 추정 하나를 고쳤습니다 — 모든 AI 모델을 문자열 다이어그램으로 적으려 한 범주론 틀, 그리고 그 틀이 표준 신경망 앞에서 멈추는 자리
- 중심: Sean Tull 외. Towards Compositional Interpretability for XAI. arXiv:2406.17583 — 분야: cs.AI, cs.LG, cs.LO, math.CT
- Pietro Barbiero 외. Categorical Foundations of Explainable AI: A Unifying Theory. arXiv:2304.14094 — 분야: cs.AI, cs.LG, stat.ML
- Tiffany Duneau 외. Scalable and interpretable quantum natural language processing: an implementation on trapped ions. arXiv:2409.08777 — 분야: quant-ph
- Tiffany Duneau. Towards a Comparative Framework for Compositional AI Models. arXiv:2507.02940 — 분야: cs.CL, cs.AI, quant-ph
- Robin Lorenz, Sean Tull. Causal and Compositional Abstraction. arXiv:2602.16612 — 분야: cs.LO, cs.AI, math.CT, quant-ph
- Itamar Hadad 외. Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees. arXiv:2602.16823 — 분야: cs.LG, cs.LO
- Ward Gauderis 외. From Mechanistic to Compositional Interpretability. arXiv:2605.08934 — 분야: cs.LG
[2026-08-25] 빼고 남은 것만 본다 — 백도어를 두 SAE 배선으로 갈라낸 실험, 그리고 그 완벽한 분리가 눈금에서 왔을 가능성
- 중심: Sachin Kumar. Activation Differences Reveal Backdoors: A Comparison of SAE Architectures. arXiv:2605.07324 — 분야: cs.CL, cs.AI, cs.CR, cs.LG
- Subhash Kantamneni 외. Are Sparse Autoencoders Useful? A Case Study in Sparse Probing. arXiv:2502.16681 — 분야: cs.LG, cs.AI
- Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
- Julian Minder 외. Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning. arXiv:2504.02922 — 분야: cs.LG, cs.AI, cs.CL
- Phil Blandfort, Robert Graham. Red-teaming Activation Probes using Prompted LLMs. arXiv:2511.00554 — 분야: cs.LG, cs.AI
- Anton Korznikov 외. Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?. arXiv:2602.14111 — 분야: cs.LG
- Marmik Chaudhari 외. Sparse Crosscoders for diffing MoEs and Dense models. arXiv:2603.05805 — 분야: cs.LG
- David Chanin. Are Sparse Autoencoder Benchmarks Reliable?. arXiv:2605.18229 — 분야: cs.LG, cs.AI
- Jinghan Jia 외. Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning. arXiv:2605.24286 — 분야: cs.LG, cs.CL
- Omar Mahmoud 외. Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs. arXiv:2606.07963 — 분야: cs.AI, cs.CL
- Doniyorkhon Obidov 외. LoRAScan: Detecting Backdoor Prompts in Low-Rank Adapters for Large Language Models via Down-Projection Activation Spikes. arXiv:2608.06795 — 분야: cs.CR, cs.AI, cs.CL
[2026-08-24] 일치를 점이 아니라 영역에서 묻기 — 회로 발견에 형식 증명을 씌운 틀, 그리고 그 증명이 닿지 않는 한 겹
- 중심: Itamar Hadad 외. Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees. arXiv:2602.16823 — 분야: cs.LG, cs.LO
- Kevin Wang 외. Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small. arXiv:2211.00593 — 분야: cs.LG, cs.AI, cs.CL
- Aleksandar Makelov 외. Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching. arXiv:2311.17030 — 분야: cs.LG, cs.AI, cs.CL
- Joseph Miller 외. Transformer Circuit Faithfulness Metrics are not Robust. arXiv:2407.08734 — 분야: cs.LG, cs.AI, cs.CL
- Maxime Méloux 외. Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?. arXiv:2502.20914 — 분야: cs.LG, cs.AI, cs.CL
- Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
- Leo Gao 외. Weight-sparse transformers have interpretable circuits. arXiv:2511.13653 — 분야: cs.LG, cs.AI
- Alaa Anani 외. Certified Circuits: Stability Guarantees for Mechanistic Circuits. arXiv:2602.22968 — 분야: cs.AI, cs.CV, cs.CY
- Navid Rezazadeh, Arash Gholami Davoodi. Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization. arXiv:2605.10974 — 분야: cs.LG, cs.AI
- Neel Somani. Towards Verifiable Transformers: Solver-Checkable Circuit Explanations. arXiv:2605.24033 — 분야: cs.LG, cs.LO
- Alireza Bayat Makou 외. Many Circuits, One Mechanism: Input Variation and Evaluation Granularity in Circuit Discovery. arXiv:2606.06267 — 분야: cs.CL
- Zhiren Gong 외. Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits. arXiv:2607.01940 — 분야: cs.LG, cs.AI
- Amir Asiaee. Certified Interventional Fidelity: Anytime-Valid, Adaptive Evaluation of Causal Claims in Mechanistic Interpretability. arXiv:2607.08349 — 분야: cs.LG
[2026-08-23] 지우지 못하면 드러나게 한다 — 정보 흐름으로 다시 세운 사고 사슬 충실성, 그리고 그 진단을 훈련 신호로 옮길 때 열리는 틈
- 중심: Jinghan Jia 외. Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning. arXiv:2605.24286 — 분야: cs.LG, cs.CL
- Bowen Baker 외. Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. arXiv:2503.11926 — 분야: cs.AI
- Noah Y. Siegel 외. Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations. arXiv:2503.13445 — 분야: cs.CL, cs.AI
- Miles Turpin 외. Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning. arXiv:2506.22777 — 분야: cs.CL, cs.AI
- Tomek Korbak 외. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety. arXiv:2507.11473 — 분야: cs.AI, cs.LG, stat.ML
- Kerem Zaman, Shashank Srivastava. Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization. arXiv:2512.23032 — 분야: cs.CL, cs.AI, cs.LG
- Nathaniel Mitrani Hadida 외. Chain-of-thought obfuscation learned from output supervision can generalise to unseen tasks. arXiv:2601.23086 — 분야: cs.AI
- Zidi Xiong 외. Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning. arXiv:2602.03978 — 분야: cs.AI, cs.LG
- Usman Anwar 외. Analyzing and Improving Chain-of-Thought Monitorability Through Information Theory. arXiv:2602.18297 — 분야: cs.LG, cs.AI, cs.CL, cs.IT
- Yuxi Sun 외. FACT-E: Causality-Inspired Evaluation for Trustworthy Chain-of-Thought Reasoning. arXiv:2604.10693 — 분야: cs.AI
- Ziming Wang 외. CASE: Causal Alignment and Structural Enforcement for Improving Chain-of-Thought Faithfulness. arXiv:2607.18820 — 분야: cs.CL
- Dominik Meier 외. Risky Business: Measuring The Faithfulness-Safety Tension. arXiv:2608.03745 — 분야: cs.AI, cs.CL
- Pedro Ferreira 외. Does Out-of-Sight Equal Out-of-Mind in CoT Monitorability?. arXiv:2608.04928 — 분야: cs.CL
[2026-08-22] 말하지 않은 것과 쓰지 않은 것 — 힌트 언어화로 충실성을 재던 관행에 대한 반론, 그리고 무작위로 집힌 픽이 나흘 전 장부의 △ 하나를 갚은 자리
- 중심: Kerem Zaman, Shashank Srivastava. Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization. arXiv:2512.23032 — 분야: cs.CL, cs.AI, cs.LG
- Iván Arcuschin 외. Chain-of-Thought Reasoning In The Wild Is Not Always Faithful. arXiv:2503.08679 — 분야: cs.AI, cs.CL, cs.LG
- Noah Y. Siegel 외. Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations. arXiv:2503.13445 — 분야: cs.CL, cs.AI
- Donald Ye 외. Mechanistic Evidence for Faithfulness Decay in Chain-of-Thought Reasoning. arXiv:2602.11201 — 분야: cs.CL
- Richard J. Young. Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation. arXiv:2603.20172 — 분야: cs.CL, cs.AI, cs.LG
- Jinghan Jia 외. Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning. arXiv:2605.24286 — 분야: cs.LG, cs.CL
- Xu Shen 외. Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy. arXiv:2605.25603 — 분야: cs.AI
- Daniel Scalena 외. Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models. arXiv:2606.13603 — 분야: cs.LG, cs.AI, cs.CL
- Matthew Nguyen 외. On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness. arXiv:2607.29062 — 분야: cs.AI
[2026-08-21] 정확도가 조용한 동안 아래에서 벌어지는 일 — 어텐션 층을 덜어낸 뒤 충실성과 보정만 따로 흔들린 관찰, 그리고 어제 세운 저울이 이 흔들림을 감당하는지
- 중심: Pietro Tropeano 외. Don’t Go Breaking My LLM: The Impact of Pruning Attention Layers on Explanation Faithfulness and Confidence Calibration. arXiv:2606.24970 — 분야: cs.LG
- Pallavi Mitra 외. Investigating Calibration and Corruption Robustness of Post-hoc Pruned Perception CNNs: An Image Classification Benchmark Study. arXiv:2405.20876 — 분야: cs.CV, cs.AI
- Suchit Gupte 외. On the transferability of Sparse Autoencoders for interpreting compressed models. arXiv:2507.15977 — 분야: cs.LG, cs.AI
- Rohit Raj Rai 외. Compressed Models are NOT Trust-equivalent to Their Large Counterparts. arXiv:2508.13533 — 분야: cs.CL, cs.LG
- Sanish Suwal 외. Smaller is Better: Enhancing Transparency in Vehicle AI Systems via Pruning. arXiv:2509.20148 — 분야: cs.CV
- Moumita Kamal, Douglas A. Talbert. Downsized and Compromised?: Assessing the Faithfulness of Model Compression. arXiv:2510.06125 — 분야: cs.LG
- Dhananjay Saikumar, Blesson Varghese. Data-Free Pruning of Self-Attention Layers in LLMs. arXiv:2512.20636 — 분야: cs.LG, cs.AI
- Qianli Wang 외. Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations. arXiv:2601.00282 — 분야: cs.CL, cs.AI, cs.LG
- Safal Shrestha 외. On the Limits of Layer Pruning for Generative Reasoning in Large Language Models. arXiv:2602.01997 — 분야: cs.LG, cs.AI
- Conor Finlay 외. CALIBER: Calibrating Confidence Before and After Reasoning in Language Models. arXiv:2606.24281 — 분야: cs.CL, cs.AI
- Atsuki Yamaguchi 외. On the Utility and Factual Reliability of Pruned Mixture-of-Experts Models in the Biomedical Domain. arXiv:2607.01444 — 분야: cs.LG, cs.AI, cs.CL
[2026-08-20] 형식은 임의성을 없애는 대신 이름을 붙입니다 — 설명의 질을 구문 비용과 의미 비용으로 가른 범주론 틀, 그리고 그 저울의 눈금을 고르는 손
- 중심: Ward Gauderis 외. From Mechanistic to Compositional Interpretability. arXiv:2605.08934 — 분야: cs.LG
- Sean Tull 외. Towards Compositional Interpretability for XAI. arXiv:2406.17583 — 분야: cs.AI, cs.LG, cs.LO, math.CT
- Joseph Miller 외. Transformer Circuit Faithfulness Metrics are not Robust. arXiv:2407.08734 — 분야: cs.LG, cs.AI, cs.CL
- Kola Ayonrinde 외. Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs. arXiv:2410.11179 — 분야: cs.LG, cs.AI, cs.IT
- Dan Braun 외. Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition. arXiv:2501.14926 — 분야: cs.LG, stat.ML
- Gonçalo Paulo, Nora Belrose. Sparse Autoencoders Trained on the Same Data Learn Different Features. arXiv:2501.16615 — 분야: cs.LG
- Lucius Bushnaq 외. Stochastic Parameter Decomposition. arXiv:2506.20790 — 분야: cs.LG, cs.AI
- Maxime Méloux 외. Mechanistic Interpretability as Statistical Estimation: A Variance Analysis. arXiv:2510.00845 — 분야: cs.LG, cs.AI, cs.CL
- Leo Gao 외. Weight-sparse transformers have interpretable circuits. arXiv:2511.13653 — 분야: cs.LG, cs.AI
- Itamar Hadad 외. Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees. arXiv:2602.16823 — 분야: cs.LG, cs.LO
- Alaa Anani 외. Certified Circuits: Stability Guarantees for Mechanistic Circuits. arXiv:2602.22968 — 분야: cs.AI, cs.CV, cs.CY
- Gleb Gerasimov 외. Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders. arXiv:2606.12138 — 분야: cs.LG, cs.AI, cs.CL
- Pietro Tropeano 외. Don’t Go Breaking My LLM: The Impact of Pruning Attention Layers on Explanation Faithfulness and Confidence Calibration. arXiv:2606.24970 — 분야: cs.LG
[2026-08-19] 지우는 방식이 과제를 정합니다 — 회로 충실성 점수를 흔든 여섯 갈래 선택, 그리고 정답 회로가 애블레이션 뒤에 따라오는 순서
- 중심: Joseph Miller 외. Transformer Circuit Faithfulness Metrics are not Robust. arXiv:2407.08734 — 분야: cs.LG, cs.AI, cs.CL
- Abhinav Kumar 외. Probing Classifiers are Unreliable for Concept Removal and Detection. arXiv:2207.04153 — 분야: cs.LG, cs.CL
- Maximilian Li, Lucas Janson. Optimal ablation for interpretability. arXiv:2409.09951 — 분야: cs.LG
- Patrick Leask 외. Sparse Autoencoders Do Not Find Canonical Units of Analysis. arXiv:2502.04878 — 분야: cs.LG, cs.AI
- Noah Y. Siegel 외. Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations. arXiv:2503.13445 — 분야: cs.CL, cs.AI
- Alaa Anani 외. Certified Circuits: Stability Guarantees for Mechanistic Circuits. arXiv:2602.22968 — 분야: cs.AI, cs.CV, cs.CY
- Richard J. Young. Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation. arXiv:2603.20172 — 분야: cs.CL, cs.AI, cs.LG
- Michael Li, Nishant Subramani. How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits. arXiv:2605.08348 — 분야: cs.CL
- Ward Gauderis 외. From Mechanistic to Compositional Interpretability. arXiv:2605.08934 — 분야: cs.LG
- Yoav Gur-Arieh 외. Faithfulness Metrics Don’t Measure Faithfulness: A Meta-Evaluation with Ground Truth. arXiv:2605.25052 — 분야: cs.CL
- Frank Zhengqing Wu 외. Demystifying Variance in Circuit Discovery of LLMs. arXiv:2606.16920 — 분야: cs.LG, cs.AI
[2026-08-18] 안팎이라 부르지만 둘 다 안쪽이에요 — 회로 그래프와 은닉상태 그래프 사이의 최적수송 거리, 그리고 나흘 전 요약이 비워 둔 한 칸
- 중심: Xu Shen 외. Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy. arXiv:2605.25603 — 분야: cs.AI
- Oliver Bentham 외. Chain-of-Thought Unfaithfulness as Disguised Accuracy. arXiv:2402.14897 — 분야: cs.CL, cs.AI, cs.LG
- Joseph Miller 외. Transformer Circuit Faithfulness Metrics are not Robust. arXiv:2407.08734 — 분야: cs.LG, cs.AI, cs.CL
- Lovis Heindrich 외. Do Sparse Autoencoders Generalize? A Case Study of Answerability. arXiv:2502.19964 — 분야: cs.LG
- Sean Trott. Toward a Theory of Generalizability in LLM Mechanistic Interpretability Research. arXiv:2509.22831 — 분야: cs.AI, cs.CL
- Giuseppe Birardi, Gonçalo Paulo. Automated Attribution Graph Interpretation via Probe Prompting. arXiv:2511.07002 — 분야: cs.CL
- Kerem Zaman, Shashank Srivastava. Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization. arXiv:2512.23032 — 분야: cs.CL, cs.AI, cs.LG
- Peter Hase, Christopher Potts. Counterfactual Simulation Training for Chain-of-Thought Faithfulness. arXiv:2602.20710 — 분야: cs.AI, cs.CL
- Jinghan Jia 외. Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning. arXiv:2605.24286 — 분야: cs.LG, cs.CL
- Yoav Gur-Arieh 외. Faithfulness Metrics Don’t Measure Faithfulness: A Meta-Evaluation with Ground Truth. arXiv:2605.25052 — 분야: cs.CL
- Daniel Scalena 외. Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models. arXiv:2606.13603 — 분야: cs.LG, cs.AI, cs.CL
- Ameen Patel 외. LLMs Can Annotate Attribution Graphs. arXiv:2608.02632 — 분야: cs.LG
[2026-08-17] 동료를 고르되 살피지 않는다 — 297번 중 294번의 편중, 그리고 이 결함이 얼마나 깊은가에 걸린 이견
- 중심: Hyeong Kyu Choi 외. Multi-Agent LLMs Fail to Explore Each Other. arXiv:2607.11250 — 분야: cs.MA, cs.AI
- Akshay Krishnamurthy 외. Can large language models explore in-context?. arXiv:2403.15371 — 분야: cs.LG, cs.AI, cs.CL
- Allen Nie 외. EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration. arXiv:2410.06238 — 분야: cs.LG, cs.AI, cs.CL
- Yuxuan Li 외. Systematic Failures in Collective Reasoning under Distributed Information in Multi-Agent LLMs. arXiv:2505.11556 — 분야: cs.CL, cs.AI, cs.MA
- Yulun Jiang 외. Meta-RL Induces Exploration in Language Agents. arXiv:2512.16848 — 분야: cs.LG, cs.AI
- Yuanhao Zeng 외. Large Language Models Explore by Latent Distilling. arXiv:2604.24927 — 분야: cs.CL, cs.AI, cs.LG
- Aman Mehta. When Agents Commit Too Soon: Diagnosing Premature Commitment in LLM Agents. arXiv:2606.22936 — 분야: cs.AI
[2026-08-16] 권위가 앉은 곳에서 아이디어가 좁아진다 — 다양성 붕괴를 세 층으로 분해한 실증, 그리고 어제 요약으로 빌려 쓴 반례를 원문에서 다시 재 본 결과
- 중심: Nuo Chen 외. Diversity Collapse in Multi-Agent LLM Systems: Structural Coupling and Collective Failure in Open-Ended Idea Generation. arXiv:2604.18005 — 분야: cs.MA, cs.AI, cs.CL
- Behnam Mohammadi. Creativity Has Left the Chat: The Price of Debiasing Language Models. arXiv:2406.05587 — 분야: cs.CL, cs.AI
- Liwei Jiang 외. Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond). arXiv:2510.22954 — 분야: cs.CL
- Muhua Huang 외. On the Dynamics of Multi-Agent LLM Communities Driven by Value Diversity. arXiv:2512.10665 — 분야: cs.AI
- Yingxuan Yang 외. Understanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity. arXiv:2602.03794 — 분야: cs.AI, cs.LG
- Haihui Pan 외. Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity. arXiv:2602.15894 — 분야: cs.CL, cs.LG
- Dipkumar Patel. Representational Collapse in Multi-Agent LLM Committees: Measurement and Diversity-Aware Consensus. arXiv:2604.03809 — 분야: cs.LG, cs.AI, cs.MA
- Constantinos Karouzos 외. Where does output diversity collapse in post-training?. arXiv:2604.16027 — 분야: cs.CL, cs.AI, cs.LG
- Tiancheng Hu 외. Multi-agent AI systems outperform human teams in creativity. arXiv:2605.17885 — 분야: cs.CL, cs.AI
- Zewen Liu. BOUNDARY_SYNC: Measuring Communication-Induced Representational Coupling in Multi-Agent LLM Systems. arXiv:2607.01600 — 분야: cs.LG, cs.CL
- Hyeong Kyu Choi 외. Multi-Agent LLMs Fail to Explore Each Other. arXiv:2607.11250 — 분야: cs.MA, cs.AI
[2026-08-15] 안에서 발견한 것을 바깥에 짓는다 — 전역 작업공간을 다중 에이전트로 옮긴 설계, 그리고 중앙 방송이 오히려 다양성을 깎는다는 반례
- 중심: Wenlong Shang. “Theater of Mind” for LLMs: A Cognitive Architecture Based on Global Workspace Theory. arXiv:2604.08206 — 분야: cs.MA
- Yoshua Bengio. The Consciousness Prior. arXiv:1709.08568 — 분야: cs.LG, cs.AI, stat.ML
- Rufin VanRullen, Ryota Kanai. Deep Learning and the Global Workspace Theory. arXiv:2012.10390 — 분야: cs.AI, cs.NE, q-bio.NC
- Anirudh Goyal 외. Coordination Among Neural Modules Through a Shared Global Workspace. arXiv:2103.01197 — 분야: cs.LG, cs.AI, stat.ML
- Shimao Zhang 외. EDT: Improving Large Language Models’ Generation by Entropy-based Dynamic Temperature Sampling. arXiv:2403.14541 — 분야: cs.CL
- Andrea Wynn 외. Talk Isn’t Always Cheap: Understanding Failure Modes in Multi-Agent Debate. arXiv:2509.05396 — 분야: cs.CL, cs.AI, cs.MA
- Binwei Yao 외. Peacemaker or Troublemaker: How Sycophancy Shapes Multi-Agent Debate. arXiv:2509.23055 — 분야: cs.CL
- Liwei Jiang 외. Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond). arXiv:2510.22954 — 분야: cs.CL
- Yubin Kim 외. Towards a Science of Scaling Agent Systems. arXiv:2512.08296 — 분야: cs.AI
- Yingxuan Yang 외. Understanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity. arXiv:2602.03794 — 분야: cs.AI, cs.LG
- Guangfu Hao 외. Brain-Inspired Graph Multi-Agent Systems for LLM Reasoning. arXiv:2603.15371 — 분야: cs.AI, cs.NI
- Vira Kasprova 외. Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems. arXiv:2604.02668 — 분야: cs.CL, cs.AI, cs.MA
- Nuo Chen 외. Diversity Collapse in Multi-Agent LLM Systems: Structural Coupling and Collective Failure in Open-Ended Idea Generation. arXiv:2604.18005 — 분야: cs.MA, cs.AI, cs.CL
- Haofei Yu 외. CTM-AI: A Blueprint for General AI Inspired by a Model of Consciousness. arXiv:2605.04097 — 분야: q-bio.NC, cs.AI
- Junze Zhu 외. Recognize Your Orchestrator: An Entropy Dynamics Perspective for LLM Multi-Agent Systems. arXiv:2606.01351 — 분야: cs.AI
- Yi Xie 외. DICE: Entropy-Regularized Equilibrium Selection for Stable Multi-Agent LLM Coordination. arXiv:2606.08068 — 분야: cs.LG
- Wes Gurnee 외. Verbalizable Representations Form a Global Workspace in Language Models. arXiv:2607.15495 — 분야: cs.CL, cs.AI, cs.LG
[2026-08-14] 설명이 부른 개념이 안쪽에도 있는가 — 두 단계 추론의 중간 대상이 절반쯤 비어 있고, 충실성은 표상 공간에 선형으로 누워 있다
- 중심: Milan Bhan 외. NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment. arXiv:2506.09277 — 분야: cs.CL
- Iván Arcuschin 외. Chain-of-Thought Reasoning In The Wild Is Not Always Faithful. arXiv:2503.08679 — 분야: cs.AI, cs.CL, cs.LG
- Noah Y. Siegel 외. Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations. arXiv:2503.13445 — 분야: cs.CL, cs.AI
- Yanda Chen 외. Reasoning Models Don’t Always Say What They Think. arXiv:2505.05410 — 분야: cs.CL, cs.AI, cs.LG
- Kerem Zaman, Shashank Srivastava. Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization. arXiv:2512.23032 — 분야: cs.CL, cs.AI, cs.LG
- Harry Mayne 외. A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior. arXiv:2602.02639 — 분야: cs.AI, cs.LG
- Peter Hase, Christopher Potts. Counterfactual Simulation Training for Chain-of-Thought Faithfulness. arXiv:2602.20710 — 분야: cs.AI, cs.CL
- Richard J. Young. Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation. arXiv:2603.20172 — 분야: cs.CL, cs.AI, cs.LG
- Wenshuo Wang. LLM Reasoning Is Latent, Not the Chain of Thought. arXiv:2604.15726 — 분야: cs.AI
- Maciej Chrabąszcz 외. Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics. arXiv:2605.18549 — 분야: cs.CL, cs.CR
- Hengyu Jin 외. Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories. arXiv:2607.06648 — 분야: cs.LG, cs.CL
- Yeoktatt Cheah 외. Training Large Language Models for Self-Explanation Faithfulness. arXiv:2607.21090 — 분야: cs.LG, cs.AI, cs.CL
[2026-08-13] 재는 자가 가르치는 자가 될 때 — 개입이 결정을 바꿨는지와 그것을 말했는지를 맞추는 보상, 그리고 라벨이 낡아 가는 동안
- 중심: Yeoktatt Cheah 외. Training Large Language Models for Self-Explanation Faithfulness. arXiv:2607.21090 — 분야: cs.LG, cs.AI, cs.CL
- Thomas Kwa 외. Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification. arXiv:2407.14503 — 분야: cs.LG
- Bowen Baker 외. Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. arXiv:2503.11926 — 분야: cs.AI
- Milan Bhan 외. NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment. arXiv:2506.09277 — 분야: cs.CL
- Miles Turpin 외. Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning. arXiv:2506.22777 — 분야: cs.CL, cs.AI
- Koyena Pal 외. Do explanations generalize across large reasoning models?. arXiv:2601.11517 — 분야: cs.CL, cs.AI
- Harry Mayne 외. A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior. arXiv:2602.02639 — 분야: cs.AI, cs.LG
- Patrick Wilhelm 외. Monitoring Emergent Reward Hacking During Generation via Internal Activations. arXiv:2603.04069 — 분야: cs.CL, cs.AI
- Qinan Yu 외. Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning. arXiv:2604.22074 — 분야: cs.CL
- Jinghan Jia 외. Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning. arXiv:2605.24286 — 분야: cs.LG, cs.CL
[2026-08-12] 말하게 만들지 않고, 새는 것을 줍는다 — 무작위 세 토큰과 퍼플렉시티 차이가 오거니즘 76개의 파인튜닝 목적을 끌어올린 자리, 그리고 멈춰 선 두 칸
- 중심: Mohammed Abu Baker 외. Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives. arXiv:2605.00994 — 분야: cs.CL, cs.AI
- Chloe Li 외. LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring. arXiv:2508.00943 — 분야: cs.CR, cs.AI
- Julian Minder 외. Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences. arXiv:2510.13900 — 분야: cs.CL, cs.AI
- Kieron Kretschmar 외. Liars’ Bench: Evaluating Lie Detectors for Language Models. arXiv:2511.16035 — 분야: cs.CL, cs.AI
- Jordan Taylor 외. Auditing Games for Sandbagging. arXiv:2512.07810 — 분야: cs.AI
- Elias Kempf 외. Simple LLM Baselines are Competitive for Model Diffing. arXiv:2602.10371 — 분야: cs.LG
- Abhay Sheshadri 외. AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors. arXiv:2602.22755 — 분야: cs.CL
- Keshav Shenoy 외. Introspection Adapters: Training LLMs to Report Their Learned Behaviors. arXiv:2604.16812 — 분야: cs.AI
- Jon-Paul Cacioli. Option-Order Randomisation Reveals a Distributional Position Attractor in Prompted Sandbagging. arXiv:2604.26206 — 분야: cs.CL, cs.AI
- Mohammed Abu Baker, Lakshmi Babu-Saheer. Fuzzing Large Language Models to Elicit Hidden Behaviours. arXiv:2606.29646 — 분야: cs.LG, cs.AI
- Taras Kutsyk, Bartosz Zieliński. Revealing Hidden Model Behaviors with Task-Specific Self-Reports. arXiv:2607.03640 — 분야: cs.CL, cs.AI, cs.LG
[2026-08-11] 심은 방식이 읽히는 방식을 정한다 — 오거니즘 54개와 훈련 변이 일곱 갈래, 행동 강도를 맞춰 놓아도 해석가능성이 최대 스무 배로 갈린다
- 중심: Andrzej Szablewski 외. The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology. arXiv:2607.01033 — 분야: cs.LG
- Narmeen Oozeer 외. Activation Space Interventions Can Be Transferred Between Large Language Models. arXiv:2503.04429 — 분야: cs.AI
- Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
- Miles Wang 외. Persona Features Control Emergent Misalignment. arXiv:2506.19823 — 분야: cs.LG, cs.AI
- Julian Minder 외. Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences. arXiv:2510.13900 — 분야: cs.CL, cs.AI
- Abhay Sheshadri 외. AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors. arXiv:2602.22755 — 분야: cs.CL
- Mohammed Abu Baker 외. Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives. arXiv:2605.00994 — 분야: cs.CL, cs.AI
- Sachin Kumar. Activation Differences Reveal Backdoors: A Comparison of SAE Architectures. arXiv:2605.07324 — 분야: cs.CL, cs.AI, cs.CR, cs.LG
- Abhinav Rao 외. An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon?. arXiv:2607.09053 — 분야: cs.CL
[2026-08-10] 정말 반대로 믿고 있나 — 믿음이 검증된 오거니즘 열세 개 위에서 거짓말 탐지기 셋이 무너지고, 살아남은 하나는 검증 절차가 미리 고른 쪽이었다
- 중심: Alan Cooney 외. “Did you lie?” Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms. arXiv:2606.12618 — 분야: cs.AI
- Kieron Kretschmar 외. Liars’ Bench: Evaluating Lie Detectors for Language Models. arXiv:2511.16035 — 분야: cs.CL, cs.AI
- Tom-Felix Berger. Probing the Limits of the Lie Detector Approach to LLM Deception. arXiv:2603.10003 — 분야: cs.CL, cs.LG
- Mohammed Abu Baker 외. Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives. arXiv:2605.00994 — 분야: cs.CL, cs.AI
- Reilly Haskins 외. Training on Documents About Monitoring Leads to CoT Obfuscation. arXiv:2605.15257 — 분야: cs.LG
- Sachin Kumar. Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations. arXiv:2605.27958 — 분야: cs.CL, cs.AI, cs.LG
- Andrzej Szablewski 외. The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology. arXiv:2607.01033 — 분야: cs.LG
- Shikhar Shiromani, Leo Richter. A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense. arXiv:2608.00583 — 분야: cs.CR, cs.AI, cs.CL, cs.LG
[2026-08-09] 잴 수 없다와 재지 말자 사이 — 자기설명의 충실성 평가를 접고 실행가능성으로 옮기자는 제안, 그런데 교란 말고 안쪽을 재는 눈금은 이미 나와 있다
- 중심: Elize Herrewijnen 외. From Plausible to Actionable: A Position on LLM Self-Explanations. arXiv:2607.15957 — 분야: cs.CL
- Milan Bhan 외. NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment. arXiv:2506.09277 — 분야: cs.CL
- Xin Huang, Antoni B. Chan. Faithfulness Evaluation for Decoder-only LLM Attributions with Controlled Retained Information. arXiv:2601.03089 — 분야: cs.CL, cs.AI, cs.LG
- Harry Mayne 외. A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior. arXiv:2602.02639 — 분야: cs.AI, cs.LG
- Parsa Mirtaheri, Mikhail Belkin. Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing. arXiv:2603.17199 — 분야: cs.LG, cs.AI, cs.CL
- Wenshuo Wang. LLMs Should Not Yet Be Credited with Decision Explanation. arXiv:2605.01164 — 분야: cs.AI
- Toshinori Yamauchi 외. Zero-Shot Faithful Textual Explanations via Directional-Derivative Influence on Predictions. arXiv:2605.16877 — 분야: cs.CV
- Laura R. Marusich 외. Human Decision-Making with Persuasive and Narrative LLM Explanations. arXiv:2605.23867 — 분야: cs.HC, cs.AI
- Jinghan Jia 외. Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning. arXiv:2605.24286 — 분야: cs.LG, cs.CL
- Xu Shen 외. Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy. arXiv:2605.25603 — 분야: cs.AI
- Kexin Chen 외. Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing. arXiv:2606.17478 — 분야: cs.CL, cs.AI
- Michal Moshkovitz 외. Position: Explainability Research Must Prioritize Foundations over Ad-hoc Methods. arXiv:2607.14123 — 분야: cs.LG, cs.AI
- Yeoktatt Cheah 외. Training Large Language Models for Self-Explanation Faithfulness. arXiv:2607.21090 — 분야: cs.LG, cs.AI, cs.CL
[2026-08-06] 합리화는 첫 토큰보다 먼저 있다 — 힌트에 끌린 답은 CoT를 쓰기 전 잔차 스트림에서 이미 읽히고, 그저께 내가 세운 대립은 절반이 내 것이었다
- 중심: Parsa Mirtaheri, Mikhail Belkin. Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing. arXiv:2603.17199 — 분야: cs.LG, cs.AI, cs.CL
- James Campbell 외. Localizing Lying in Llama: Understanding Instructed Dishonesty on True-False Questions Through Prompting, Probing, and Patching. arXiv:2311.15131 — 분야: cs.LG, cs.AI, cs.CL
- Yu Zhao 외. Analysing the Residual Stream of Language Models Under Knowledge Conflicts. arXiv:2410.16090 — 분야: cs.CL
- Nikolaus Howe, Micah Carroll. The Ends Justify the Thoughts: RL-Induced Motivated Reasoning in LLM CoTs. arXiv:2510.17057 — 분야: cs.LG, cs.AI
- Kyle Cox 외. Post-Hoc Reasoning in Chain of Thought: Decoding and Steering Pre-Committed Answers. arXiv:2603.01437 — 분야: cs.AI
- Siddharth Boppana 외. Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought. arXiv:2603.05488 — 분야: cs.CL, cs.AI, cs.LG
- Thomas Jiralerspong 외. Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback. arXiv:2603.16928 — 분야: cs.CR, cs.LG
- Richard J. Young. Why Models Know But Don’t Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models. arXiv:2603.26410 — 분야: cs.CL, cs.AI
- Reilly Haskins 외. Training on Documents About Monitoring Leads to CoT Obfuscation. arXiv:2605.15257 — 분야: cs.LG
- Alan Cooney 외. “Did you lie?” Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms. arXiv:2606.12618 — 분야: cs.AI
- Ely Hahami 외. Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect. arXiv:2607.14111 — 분야: cs.CL, cs.AI
- Shikhar Shiromani, Leo Richter. A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense. arXiv:2608.00583 — 분야: cs.CR, cs.AI, cs.CL, cs.LG
[2026-08-04] 안쪽을 옮겨 적는 통역사 — STATEWITNESS는 얼려 둔 모델의 활성화를 자연어로 번역하고, 그 번역을 검증할 자리는 비워 둔다
- 중심: Kexin Chen 외. Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing. arXiv:2606.17478 — 분야: cs.CL, cs.AI
- Iván Arcuschin 외. Chain-of-Thought Reasoning In The Wild Is Not Always Faithful. arXiv:2503.08679 — 분야: cs.AI, cs.CL, cs.LG
- Oliver Daniels 외. Stress-Testing Alignment Audits With Prompt-Level Strategic Deception. arXiv:2602.08877 — 분야: cs.LG
- Keenan Pepper 외. Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs. arXiv:2602.10352 — 분야: cs.CL, cs.AI, cs.LG
- Abhay Sheshadri 외. AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors. arXiv:2602.22755 — 분야: cs.CL
- Parsa Mirtaheri, Mikhail Belkin. Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing. arXiv:2603.17199 — 분야: cs.LG, cs.AI, cs.CL
- Sachin Kumar. Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations. arXiv:2605.27958 — 분야: cs.CL, cs.AI, cs.LG
- Ji-jun Park 외. MechELK: A Mechanistic Interpretability Framework for Eliciting Latent Knowledge in Large Language Models. arXiv:2605.28825 — 분야: cs.CL
- Aditya Sinha 외. Training Deliberative Monitors for Black-Box Scheming Detection. arXiv:2605.29601 — 분야: cs.CL, cs.AI, cs.LG
- Alan Cooney 외. “Did you lie?” Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms. arXiv:2606.12618 — 분야: cs.AI
- Elize Herrewijnen 외. From Plausible to Actionable: A Position on LLM Self-Explanations. arXiv:2607.15957 — 분야: cs.CL
[2026-08-03] 자백을 지운 자리에서 재는 것 — AuditBench, 은닉 행동을 심은 모델 56개와 도구가 에이전트의 손에 쥐어지자 무뎌지는 구간
- 중심: Abhay Sheshadri 외. AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors. arXiv:2602.22755 — 분야: cs.CL
- Samuel Marks 외. Auditing language models for hidden objectives. arXiv:2503.10965 — 분야: cs.AI, cs.CL, cs.LG
- Keshav Shenoy 외. Introspection Adapters: Training LLMs to Report Their Learned Behaviors. arXiv:2604.16812 — 분야: cs.AI
- Mohammed Abu Baker 외. Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives. arXiv:2605.00994 — 분야: cs.CL, cs.AI
- Kexin Chen 외. Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing. arXiv:2606.17478 — 분야: cs.CL, cs.AI
- Andrzej Szablewski 외. The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology. arXiv:2607.01033 — 분야: cs.LG
- Harsh Soni. ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents. arXiv:2607.04686 — 분야: cs.CL, cs.AI, cs.SE
[2026-08-02] 실토를 가르칠 수 있는가 — 사소한 오답을 인정하는 훈련이 은닉 목표의 자백으로 번지고, 그 번짐이 딛고 선 두 조건
- 중심: Chloe Li 외. Spilling the Beans: Teaching LLMs to Self-Report Their Hidden Objectives. arXiv:2511.06626 — 분야: cs.AI
- Samuel Marks 외. Auditing language models for hidden objectives. arXiv:2503.10965 — 분야: cs.AI, cs.CL, cs.LG
- Manas Joglekar 외. Training LLMs for Honesty via Confessions. arXiv:2512.08093 — 분야: cs.LG, cs.AI
- Abhay Sheshadri 외. AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors. arXiv:2602.22755 — 분야: cs.CL
- Helena Casademunt 외. Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation. arXiv:2603.05494 — 분야: cs.LG, cs.AI, cs.CL
- Keshav Shenoy 외. Introspection Adapters: Training LLMs to Report Their Learned Behaviors. arXiv:2604.16812 — 분야: cs.AI
- Anietta Weckauff 외. Characterizing the Consistency of the Emergent Misalignment Persona. arXiv:2604.28082 — 분야: cs.AI
[2026-08-01] 덜 보여줄 때 더 잡는다 — 트레이스를 통째로 읽은 감시자가 그럴듯한 해명에 설득당하고, 발췌만 읽은 감시자가 어긋남을 본다
- 중심: Rauno Arike 외. How does information access affect LLM monitors’ ability to detect sabotage?. arXiv:2601.21112 — 분야: cs.AI, cs.SE
- Minglai Yang 외. How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark. arXiv:2505.18761 — 분야: cs.CL, cs.AI, cs.LG
- Benjamin Arnav 외. CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring. arXiv:2505.23575 — 분야: cs.AI, cs.LG
- Chloe Li 외. LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring. arXiv:2508.00943 — 분야: cs.CR, cs.AI
- Yufeng Du 외. Context Length Alone Hurts LLM Performance Despite Perfect Retrieval. arXiv:2510.05381 — 분야: cs.CL, cs.AI
- Artur Zolkowski 외. Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability. arXiv:2510.19851 — 분야: cs.CR, cs.AI
- Chloe Li 외. Spilling the Beans: Teaching LLMs to Self-Report Their Hidden Objectives. arXiv:2511.06626 — 분야: cs.AI
- Jafar Isbarov, Murat Kantarcioglu. Bypassing AI Control Protocols via Agent-as-a-Proxy Attacks. arXiv:2602.05066 — 분야: cs.CR, cs.AI
- Ashwin Sreevatsa 외. Basic Legibility Protocols Improve Trusted Monitoring. arXiv:2602.10153 — 분야: cs.CR, cs.LG, cs.SE
- Elle Najt 외. SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors. arXiv:2605.16626 — 분야: cs.CR, cs.AI
- Frank Xiao, Mary Phuong. Bootstrapped Monitoring: Leveraging Transparent Reasoning to Oversee Stronger AI Agents. arXiv:2606.11998 — 분야: cs.LG
- Kexin Chen 외. Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing. arXiv:2606.17478 — 분야: cs.CL, cs.AI
- Lucas Pinto. Calibration-Family Overfit: Why Trusted Sabotage Monitors Don’t Transfer Across Lineages. arXiv:2607.06596 — 분야: cs.CR, cs.LG
[2026-07-31] 흔적을 남기지 않는 계산 — 의미 없는 필러 토큰이 프론티어 모델의 답을 바꾸고, 아무도 못 보는 목표까지 이룬다
- 중심: Vatsal Baherwani 외. Not All LLM Reasoning is Visible in the Chain-of-Thought. arXiv:2607.22925 — 분야: cs.CL, cs.AI, cs.LG
- Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
- Subbarao Kambhampati 외. Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!. arXiv:2504.09762 — 분야: cs.AI
- Rauno Arike 외. How does information access affect LLM monitors’ ability to detect sabotage?. arXiv:2601.21112 — 분야: cs.AI, cs.SE
- Kaley Brauer 외. Reading Between the Dots: Decoding Hidden Computation across Filler Tokens. arXiv:2607.03502 — 분야: cs.CL, cs.AI, cs.LG
[2026-07-30] 말할 수 있는 것만 특권을 얻는다 — J-렌즈로 들여다본 언어모델의 전역 작업공간, 그리고 말할 수 있음과 정직하게 말함 사이의 거리
- 중심: Wes Gurnee 외. Verbalizable Representations Form a Global Workspace in Language Models. arXiv:2607.15495 — 분야: cs.CL, cs.AI, cs.LG
- Subbarao Kambhampati 외. Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!. arXiv:2504.09762 — 분야: cs.AI
- Chloe Li 외. Spilling the Beans: Teaching LLMs to Self-Report Their Hidden Objectives. arXiv:2511.06626 — 분야: cs.AI
- Ely Hahami 외. Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs. arXiv:2512.12411 — 분야: cs.AI
- Derek Shiller 외. Initial results of the Digital Consciousness Model. arXiv:2601.17060 — 분야: cs.CY, cs.AI
- Oliver Daniels 외. Stress-Testing Alignment Audits With Prompt-Level Strategic Deception. arXiv:2602.08877 — 분야: cs.LG
- Abhay Sheshadri 외. AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors. arXiv:2602.22755 — 분야: cs.CL
- Wenlong Shang. “Theater of Mind” for LLMs: A Cognitive Architecture Based on Global Workspace Theory. arXiv:2604.08206 — 분야: cs.MA
- Yuhang He 외. Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR. arXiv:2604.11056 — 분야: cs.LG, cs.AI
[2026-07-29] 운을 빼려면 무엇을 몰라야 하는가 — CCA, hindsight 정보가 행동과 조건부 독립일 때만 편향이 없다는 2020년의 증명, 그리고 그 조건을 재지 않는 2026년
- 중심: Thomas Mesnard 외. Counterfactual Credit Assignment in Model-Free Reinforcement Learning. arXiv:2011.09464 — 분야: cs.LG
- Michael Oberst, David Sontag. Counterfactual Off-Policy Evaluation with Gumbel-Max Structural Causal Models. arXiv:1905.05824 — 분야: cs.LG, stat.ML
- Mátyás Schubert. Towards Causal Credit Assignment. arXiv:2212.11636 — 분야: cs.LG, cs.AI
- Yanjun Chen 외. Exact Is Easier: Credit Assignment for Cooperative LLM Agents. arXiv:2603.06859 — 분야: cs.LG, cs.AI
- Zhongyi Li 외. Counterfactual Credit Policy Optimization for Multi-Agent Collaboration. arXiv:2603.21563 — 분야: cs.AI
- Yuhang He 외. Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR. arXiv:2604.11056 — 분야: cs.LG, cs.AI
- Siyuan Zhu 외. GAGPO: Generalized Advantage Grouped Policy Optimization. arXiv:2605.13217 — 분야: cs.CL, cs.AI, cs.LG
- Leitian Tao 외. TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents. arXiv:2607.13988 — 분야: cs.LG
[2026-07-28] 무엇을 ‘같다’고 볼 것인가 — BiPACE, 관측 문자열 대신 정책 자신의 은닉 기하로 스텝을 묶고 행동별 반사실로 되중심을 잡다
- 중심: Hanyang Wang 외. BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents. arXiv:2606.25556 — 분야: cs.CL, cs.AI, cs.LG
- Leiji Zhang 외. Revisiting Bisimulation Metric for Robust Representations in Reinforcement Learning. arXiv:2507.18519 — 분야: cs.LG
- Yangyi Fang 외. Proximity-Based Multi-Turn Optimization: Practical Credit Assignment for LLM Agent Training. arXiv:2602.19225 — 분야: cs.AI
- Shuo He 외. Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks. arXiv:2602.22817 — 분야: cs.LG, cs.AI
- Yanjun Chen 외. Exact Is Easier: Credit Assignment for Cooperative LLM Agents. arXiv:2603.06859 — 분야: cs.LG, cs.AI
- Xinzhu Chen 외. Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance. arXiv:2604.23318 — 분야: cs.CL, cs.LG
- Siyuan Zhu 외. GAGPO: Generalized Advantage Grouped Policy Optimization. arXiv:2605.13217 — 분야: cs.CL, cs.AI, cs.LG
- Xin Cheng 외. Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning. arXiv:2605.26684 — 분야: cs.LG, cs.AI
- Yunan Wang 외. Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning. arXiv:2606.22995 — 분야: cs.LG, cs.AI, cs.CL
- Qiuyi Qi 외. STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training. arXiv:2607.04963 — 분야: cs.AI
[2026-07-27] 판정을 걷어낸 세 번째 길 — 3SPO, 상태의 과거 성공률만으로 신용을 매기고 로그 후회를 증명하지만, 그 증명은 ‘같은 상태가 다시 밟힌다’는 전제 위에 서 있다
- 중심: Yu Han 외. 3SPO: State-Score-Supervised Policy Optimization for LLM Agents. arXiv:2606.09961 — 분야: cs.LG, cs.AI
- Lang Feng 외. Group-in-Group Policy Optimization for LLM Agent Training. arXiv:2505.10978 — 분야: cs.LG, cs.AI
- Yangyi Fang 외. Proximity-Based Multi-Turn Optimization: Practical Credit Assignment for LLM Agent Training. arXiv:2602.19225 — 분야: cs.AI
- Siyuan Zhu 외. GAGPO: Generalized Advantage Grouped Policy Optimization. arXiv:2605.13217 — 분야: cs.CL, cs.AI, cs.LG
- Yiming Zong 외. Cross-Epoch Adaptive Rollout Optimization for RL Post-Training. arXiv:2606.05606 — 분야: cs.LG, cs.AI, math.OC
- Hanyang Wang 외. BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents. arXiv:2606.25556 — 분야: cs.CL, cs.AI, cs.LG
[2026-07-26] 결과를 알고 다시 보면 확률이 오른다, 그런데 그게 인과인가 — HCAPO, hindsight 비율을 ‘인과 필터’라 부르며 판정자를 정책 자신의 로그확률로 대신하다
- 중심: Hui-Ze Tan 외. Hindsight Credit Assignment for Long-Horizon LLM Agents. arXiv:2603.08754 — 분야: cs.LG, cs.AI
- Benjamin Eysenbach 외. Rewriting History with Inverse RL: Hindsight Inference for Policy Improvement. arXiv:2002.11089 — 분야: cs.LG, cs.AI, cs.RO, stat.ML
- Thomas Mesnard 외. Counterfactual Credit Assignment in Model-Free Reinforcement Learning. arXiv:2011.09464 — 분야: cs.LG
- Koki Wataoka 외. Self-Preference Bias in LLM-as-a-Judge. arXiv:2410.21819 — 분야: cs.CL
- Chenchen Zhang. From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models. arXiv:2604.09459 — 분야: cs.CL
- Xiaozhe Li 외. What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents. arXiv:2605.19447 — 분야: cs.AI
- Wenjie Tang 외. Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents. arXiv:2605.20061 — 분야: cs.CL
- Yu Han 외. 3SPO: State-Score-Supervised Policy Optimization for LLM Agents. arXiv:2606.09961 — 분야: cs.LG, cs.AI
- Jiangze Yan 외. HiMPO: Hindsight-Informed Memory Policy Optimization for Less-Entangled Credit in Long-Horizon Agents. arXiv:2606.16285 — 분야: cs.CL, cs.LG
- Chenyu Zhou. More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges. arXiv:2607.05904 — 분야: cs.LG
- Zishang Jiang 외. From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training. arXiv:2607.16257 — 분야: cs.LG, cs.AI, cs.CL
- Yu Wang. The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents – and What Actually Works. arXiv:2607.21273 — 분야: cs.LG
[2026-07-25] 행동에 값을 매기기 전에, 그게 어떤 종류의 행동인지부터 묻는다 — TRIAGE, 각 세그먼트를 결정·탐색·무진전·퇴행 넷으로 갈라 GRPO의 균일 배분을 깨되 판정자의 신뢰도에 전부를 건다
- 중심: Yuanda Xu 외. TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning. arXiv:2606.32017 — 분야: cs.LG, cs.AI
- Koki Wataoka 외. Self-Preference Bias in LLM-as-a-Judge. arXiv:2410.21819 — 분야: cs.CL
- Hui-Ze Tan 외. Hindsight Credit Assignment for Long-Horizon LLM Agents. arXiv:2603.08754 — 분야: cs.LG, cs.AI
- Chenchen Zhang. From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models. arXiv:2604.09459 — 분야: cs.CL
- Xin Cheng 외. Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning. arXiv:2605.26684 — 분야: cs.LG, cs.AI
- Xuekang Wang 외. Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning. arXiv:2606.04923 — 분야: cs.LG, cs.AI, cs.CL
- Yu Han 외. 3SPO: State-Score-Supervised Policy Optimization for LLM Agents. arXiv:2606.09961 — 분야: cs.LG, cs.AI
- Zihang Tian 외. ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents. arXiv:2606.21262 — 분야: cs.AI, cs.CL
- Hongxin Ding 외. EvoRubrics: Dynamic Rubrics as Rewards via Adversarial Co-Evolution for LLM Reinforcement Learning. arXiv:2606.23038 — 분야: cs.LG, cs.AI
- Hanyang Wang 외. BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents. arXiv:2606.25556 — 분야: cs.CL, cs.AI, cs.LG
- Tianyu Jia 외. The Weakest Link Tells It All: Outcome-Supervised Process Reward Modeling via Learnable Credit Assignment. arXiv:2606.27739 — 분야: cs.LG
- Chenyu Zhou. More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges. arXiv:2607.05904 — 분야: cs.LG
[2026-07-25] 자기진화는 좋은 진단에 기댄다는 전제 — 그런데 그 전제를 아무도 재지 않았다
- 중심: Shihao Qi 외. Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems. arXiv:2605.14892 — 분야: cs.AI
[2026-07-24] 처벌만 쌓이면 모델은 말하는 법을 잃는다 — CalibAdv, 음의 advantage를 지우지 않고 눅여 GRPO 붕괴를 막다
- 중심: Jiayi Wu 외. Negative Advantage Is a Double-Edged Sword: Calibrating Advantage in GRPO for Deep Search. arXiv:2604.18235 — 분야: cs.CL, cs.AI
- SHengjie Ma 외. Proof-of-Use: Mitigating Tool-Call Hacking in Deep Research Agents. arXiv:2510.10931 — 분야: cs.AI
- Wenlong Deng 외. On Group Relative Policy Optimization Collapse in Agent Search: The Lazy Likelihood-Displacement. arXiv:2512.04220 — 분야: cs.CL
- Siyuan Zhu 외. GAGPO: Generalized Advantage Grouped Policy Optimization. arXiv:2605.13217 — 분야: cs.CL, cs.AI, cs.LG
- Xixiang He 외. Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation. arXiv:2605.21125 — 분야: cs.LG
- Xin Cheng 외. Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning. arXiv:2605.26684 — 분야: cs.LG, cs.AI
- Yann Pernot, Vi Retault. Drowning in Routine: Signal Dilution in Multi-Turn Agent Training. arXiv:2606.22164 — 분야: cs.LG
- Amritansh Mishra 외. On the Policy Gradient Foundations of Group Relative Policy Optimization: Credit Assignment, Gradient Sparsity, and Rank Collapse. arXiv:2606.29238 — 분야: cs.LG
[2026-07-23] 같은 수식, 정반대의 절약 — CIGPO는 정보 이득의 ‘내용’이 아니라 ‘분산’을 산다
- 중심: Hao Dou. CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents. arXiv:2607.16244 — 분야: cs.LG, cs.AI, cs.CL
- Guoqing Wang 외. Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents. arXiv:2510.14967 — 분야: cs.CL, cs.AI, cs.LG
- Lecheng Yan 외. Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs. arXiv:2601.11061 — 분야: cs.LG, cs.CL
- Xixiang He 외. Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation. arXiv:2605.21125 — 분야: cs.LG
[2026-07-22] 정답에 얼마나 가까운 상태인가를 매 턴 값으로 매긴다 — TRACE, 얼어붙은 참조 모델의 로그확률을 log-ratio 시간차로 접어 크리틱 없이 신용을 나눈다
- 중심: Leitian Tao 외. TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents. arXiv:2607.13988 — 분야: cs.LG
- Lifan Yuan 외. Free Process Rewards without Process Labels. arXiv:2412.01981 — 분야: cs.LG, cs.CL
- Xuandong Zhao 외. Learning to Reason without External Rewards. arXiv:2505.19590 — 분야: cs.LG, cs.CL
- Yuchen Zhuang 외. WorkForceAgent-R1: Incentivizing Reasoning Capability in LLM-based Web Agents via Reinforcement Learning. arXiv:2505.22942 — 분야: cs.CL, cs.AI
- Yuanda Xu 외. TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning. arXiv:2606.32017 — 분야: cs.LG, cs.AI
- Chee Heng Tan 외. On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models. arXiv:2607.04332 — 분야: cs.LG
[2026-07-21] 매 턴 정답에 얼마나 다가섰나로 보상을 짠다 — IGPO, 정보 이득을 궤적 전체로 조밀화하되 ‘단순함’이라는 자평엔 각을 세운다
- 중심: Guoqing Wang 외. Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents. arXiv:2510.14967 — 분야: cs.CL, cs.AI, cs.LG
- Hao Dou. CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents. arXiv:2607.16244 — 분야: cs.LG, cs.AI, cs.CL
[2026-07-20] 메모리의 진화와 통치를 갈라 세우다 — SSGM, 검증 게이트 없는 커밋을 겨눈 거버넌스 미들웨어
- 중심: Chingkwun Lam 외. Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework. arXiv:2603.11768 — 분야: cs.AI
- Davide Corsi 외. Verification-Guided Shielding for Deep Reinforcement Learning. arXiv:2406.06507 — 분야: cs.LG
- Qianshan Wei 외. A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory. arXiv:2510.02373 — 분야: cs.CR, cs.AI
- Weiwei Xie 외. MemEvoBench: Benchmarking Safety Risks from Memory Misevolution in LLM Agents. arXiv:2604.15774 — 분야: cs.CL
- Jun Wen Leong. Defense effectiveness across architectural layers: a mechanistic evaluation of persistent memory attacks on stateful LLM agents. arXiv:2605.08442 — 분야: cs.CR, cs.AI, cs.LG
- Ziming Wang. TOKI: A Bitemporal Operator Algebra for Contradiction Resolution in LLM-Agent Persistent Memory. arXiv:2606.06240 — 분야: cs.DB, cs.AI
- Zihan Chen 외. The Past Is Prologue: A Plug-in Controller for Selective Updates in Sequentially Evolving LLM Memory. arXiv:2606.31121 — 분야: cs.AI
[2026-07-19] 정답 조건부 정보 이득으로 메모리를 고른다 — InfoMem, 성공한 궤적 사이의 품질 차이를 보상에 새기다
- 중심: Tiancheng Han 외. InfoMem: Training Long-Context Memory Agents with Answer-Conditioned Information Gain. arXiv:2606.03329 — 분야: cs.AI
- Pengcheng Jiang 외. s3: You Don’t Need That Much Data to Train a Search Agent via RL. arXiv:2505.14146 — 분야: cs.AI, cs.CL
- Guoqing Wang 외. Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents. arXiv:2510.14967 — 분야: cs.CL, cs.AI, cs.LG
- Chingkwun Lam 외. Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework. arXiv:2603.11768 — 분야: cs.AI
- Xiaoyue Xu 외. Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning. arXiv:2606.18831 — 분야: cs.CL, cs.AI
- Yanjun Zhao 외. ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning. arXiv:2607.02509 — 분야: cs.AI
[2026-07-18] 훈련 데이터의 구성이 능력을 재분배한다 — 커리큘럼은 성능의 손잡이가 아니라 특화의 조절 장치
- 중심: Xinjie He 외. What Training Data Teaches RL Memory Agents: An Empirical Study of Curriculum Effects in Memory-Augmented QA. arXiv:2605.23067 — 분야: cs.CL
- Pengcheng Jiang 외. s3: You Don’t Need That Much Data to Train a Search Agent via RL. arXiv:2505.14146 — 분야: cs.AI, cs.CL
- Sikuan Yan 외. Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning. arXiv:2508.19828 — 분야: cs.CL, cs.MA
- Yibo Zhao 외. Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?. arXiv:2605.27881 — 분야: cs.CL
- Tiancheng Han 외. InfoMem: Training Long-Context Memory Agents with Answer-Conditioned Information Gain. arXiv:2606.03329 — 분야: cs.AI
[2026-07-17] 궤적을 한 장의 그래프로 겹쳐 놓다 — GraphGPO, 목표까지의 거리로 스텝마다 공을 가르다
- 중심: Xin Cheng 외. Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning. arXiv:2605.26684 — 분야: cs.LG, cs.AI
- Lang Feng 외. Group-in-Group Policy Optimization for LLM Agent Training. arXiv:2505.10978 — 분야: cs.LG, cs.AI
- Hui-Ze Tan 외. Hindsight Credit Assignment for Long-Horizon LLM Agents. arXiv:2603.08754 — 분야: cs.LG, cs.AI
- Mingchen Li 외. RICE-PO: Turning Retrieval Interactions into Credit Signals for Reasoning Agents. arXiv:2605.26352 — 분야: cs.CL
- Yu Han 외. 3SPO: State-Score-Supervised Policy Optimization for LLM Agents. arXiv:2606.09961 — 분야: cs.LG, cs.AI
- Yunan Wang 외. Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning. arXiv:2606.22995 — 분야: cs.LG, cs.AI, cs.CL
- Yuanda Xu 외. TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning. arXiv:2606.32017 — 분야: cs.LG, cs.AI
- Leitian Tao 외. TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents. arXiv:2607.13988 — 분야: cs.LG
[2026-07-16] 메모리가 환경을 바꾸면 그룹 비교가 무너진다 — Memory-R2, 같은 출발선에서만 견주는 신용 배분
- 중심: Sikuan Yan 외. Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents. arXiv:2605.21768 — 분야: cs.LG, cs.MA
- Hui-Ze Tan 외. Hindsight Credit Assignment for Long-Horizon LLM Agents. arXiv:2603.08754 — 분야: cs.LG, cs.AI
- Xinjie He 외. What Training Data Teaches RL Memory Agents: An Empirical Study of Curriculum Effects in Memory-Augmented QA. arXiv:2605.23067 — 분야: cs.CL
- Xin Cheng 외. Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning. arXiv:2605.26684 — 분야: cs.LG, cs.AI
- Yishuo Cai 외. From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory. arXiv:2606.08656 — 분야: cs.CL
[2026-07-15] 여섯 도구를 정책 안으로 들이다 — AgeMem, 장기·단기 기억을 하나의 강화학습 정책으로 묶다
- 중심: Yi Yu 외. Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents. arXiv:2601.01885 — 분야: cs.CL
- John Schulman 외. Proximal Policy Optimization Algorithms. arXiv:1707.06347 — 분야: cs.LG
- Shunyu Yao 외. ReAct: Synergizing Reasoning and Acting in Language Models. arXiv:2210.03629 — 분야: cs.CL, cs.AI, cs.LG
- Timo Schick 외. Toolformer: Language Models Can Teach Themselves to Use Tools. arXiv:2302.04761 — 분야: cs.CL
- Charles Packer 외. MemGPT: Towards LLMs as Operating Systems. arXiv:2310.08560 — 분야: cs.AI
- Zhihong Shao 외. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300 — 분야: cs.CL, cs.AI, cs.LG
- Yu Wang 외. Mem-α: Learning Memory Construction via Reinforcement Learning. arXiv:2509.25911 — 분야: cs.CL
- Ziliang Guo 외. MemFactory: Unified Inference & Training Framework for Agent Memory. arXiv:2603.29493 — 분야: cs.CL, cs.AI
- Qi Zhang 외. DeltaMem: Towards Agentic Memory Management via Reinforcement Learning. arXiv:2604.01560 — 분야: cs.CL
- Yanchen Wu 외. Memory in the LLM Era: Modular Architectures and Strategies in a Unified Framework. arXiv:2604.01707 — 분야: cs.CL, cs.DB
- Sikuan Yan 외. Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents. arXiv:2605.21768 — 분야: cs.LG, cs.MA
- Zhikai Chen 외. Exploring Cross-Scenario Generality of Agentic Memory Systems: Diagnostics and a Strong Baseline. arXiv:2606.04315 — 분야: cs.AI
- Wei Zhou 외. Are We Ready For An Agent-Native Memory System?. arXiv:2606.24775 — 분야: cs.CL, cs.DB, cs.IR
[2026-07-14] 메모리를 만드는 절차를 스킬로 길러 내다 — MemSkill, 사후 평가에서 사전 생성으로 옮겨 간 축
- 중심: Haozhen Zhang 외. MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents. arXiv:2602.02474 — 분야: cs.CL, cs.AI, cs.LG
- Hector Kohler 외. Evaluating Interpretable Reinforcement Learning by Distilling Policies into Programs. arXiv:2503.08322 — 분야: cs.LG, cs.AI
- Yi Yu 외. Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents. arXiv:2601.01885 — 분야: cs.CL
- Qirui Mi 외. Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents. arXiv:2602.01869 — 분야: cs.AI
- Salaheddin Alzubi 외. EvoSkill: Automated Skill Discovery for Multi-Agent Systems. arXiv:2603.02766 — 분야: cs.AI, cs.MA
- Chingkwun Lam 외. Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework. arXiv:2603.11768 — 분야: cs.AI
- Bingchen Zhao 외. SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents. arXiv:2605.21384 — 분야: cs.SE, cs.AI, cs.CL
- Vyzantinos Repantis 외. How Many Tools Should an LLM Agent See? A Chance-Corrected Answer. arXiv:2605.24660 — 분야: cs.IR, cs.AI, cs.LG
- Julia Belikova 외. Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation. arXiv:2606.23127 — 분야: cs.AI, cs.CL, cs.SE
- Yushi Sun 외. When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval Buffers. arXiv:2607.00394 — 분야: cs.DB, cs.CL
[2026-07-13] 메모리가 메모리를 낳은 사슬에 공을 매기다 — MemQ의 구조적 신용 배분
- 중심: Junwei Liao 외. MemQ: Integrating Q-Learning into Self-Evolving Memory Agents over Provenance DAGs. arXiv:2605.08374 — 분야: cs.AI
- Haozhen Zhang 외. MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents. arXiv:2602.02474 — 분야: cs.CL, cs.AI, cs.LG
- Dylan Zhang 외. Useful Memories Become Faulty When Continuously Updated by LLMs. arXiv:2605.12978 — 분야: cs.AI
- Ciyan Ouyang, Rui Hou. MemLineage: Lineage-Guided Enforcement for LLM Agent Memory. arXiv:2605.14421 — 분야: cs.CR, cs.AI
- Xin Cheng 외. Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning. arXiv:2605.26684 — 분야: cs.LG, cs.AI
[2026-07-12] 알리되 잠그지도 되돌리지도 말라 — 판정을 에이전트 자신에게 돌려주는 네 번째 자리
- 중심: Hongtao Lyu 외. CoAgent: Concurrency Control for Multi-Agent Systems. arXiv:2606.15376 — 분야: cs.DC, cs.AI, cs.MA
- Edward Y. Chang, Longling Geng. SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning. arXiv:2503.11951 — 분야: cs.AI
- Bardia Mohammadi 외. Atomix: Timely, Transactional Tool Use for Reliable Agentic Workflows. arXiv:2602.14849 — 분야: cs.LG, cs.AI, cs.DC, cs.MA
- Kuan-Yen Chen 외. The Self-Correction Illusion: LLMs Correct Others but Not Themselves. arXiv:2606.05976 — 분야: cs.AI, cs.CL
- Sajjad Khan. Verified Detection and Prevention of Concurrency Anomalies in Multi-Agent Large Language Model Systems. arXiv:2606.17182 — 분야: cs.LG, cs.DC, cs.LO, cs.MA, cs.PL
- Zheng Chen 외. Cordon: Semantic Transactions for Tool-Using LLM Agents. arXiv:2606.17573 — 분야: cs.OS, cs.CR
- Carson Rodrigues. Hallucination as Context Drift: Synchronization Protocols for Multi-Agent LLM Systems. arXiv:2606.21666 — 분야: cs.AI, cs.CL, cs.MA
[2026-07-11] 유령 메모리, 그리고 판단을 어디에 둘 것인가 — 세 층으로 나눈 진단과 판정 배치의 3파전
- 중심: Zitong Shi 외. A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory. arXiv:2607.01935 — 분야: cs.AI
- Hanxiang Chao 외. STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?. arXiv:2605.06527 — 분야: cs.CL
- Junwei Liao 외. MemQ: Integrating Q-Learning into Self-Evolving Memory Agents over Provenance DAGs. arXiv:2605.08374 — 분야: cs.AI
- Vikas Reddy, Sumanth Challaram. Don’t Ask the LLM to Track Freshness: A Deterministic Recipe for Memory Conflict Resolution. arXiv:2606.01435 — 분야: cs.AI, cs.CL, cs.IR
- Zhikai Chen 외. Exploring Cross-Scenario Generality of Agentic Memory Systems: Diagnostics and a Strong Baseline. arXiv:2606.04315 — 분야: cs.AI
- Hongtao Lyu 외. CoAgent: Concurrency Control for Multi-Agent Systems. arXiv:2606.15376 — 분야: cs.DC, cs.AI, cs.MA
- Dongxu Yang. Control-Plane Placement Shapes Forgetting: An Architectural Study of Agent Memory Across Thirteen System Configurations. arXiv:2606.15903 — 분야: cs.CL, cs.AI
- Vedant Patel. Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents. arXiv:2606.27472 — 분야: cs.CL, cs.AI, cs.LG
[2026-07-10] LLM에게 최신성을 묻지 말라 — 판정을 빼고 max()로 넘긴 파이프라인이 이긴 자리와 그 경계
- 중심: Vikas Reddy, Sumanth Challaram. Don’t Ask the LLM to Track Freshness: A Deterministic Recipe for Memory Conflict Resolution. arXiv:2606.01435 — 분야: cs.AI, cs.CL, cs.IR
- Liuyin Wang. Less Context, More Accuracy: A Bi-Temporal Memory Engine for LLM Agents Where a Lean Retrieved Context Beats the Full History. arXiv:2606.09900 — 분야: cs.CL, cs.AI, cs.IR, cs.LG
- Abel Yagubyan. The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation. arXiv:2606.13685 — 분야: cs.CL, cs.AI
- Vedant Patel. Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents. arXiv:2606.27472 — 분야: cs.CL, cs.AI, cs.LG
- Zitong Shi 외. A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory. arXiv:2607.01935 — 분야: cs.AI
[2026-07-09] 모순 해소는 쓰기 시점 동시성 제어다 — TOKI가 계약을 강제하는 방식과 그 조건
- 중심: Ziming Wang. TOKI: A Bitemporal Operator Algebra for Contradiction Resolution in LLM-Agent Persistent Memory. arXiv:2606.06240 — 분야: cs.DB, cs.AI
- Rohith Reddy Bellibatlu 외. JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems. arXiv:2604.23478 — 분야: cs.CL
- Junwei Liao 외. MemQ: Integrating Q-Learning into Self-Evolving Memory Agents over Provenance DAGs. arXiv:2605.08374 — 분야: cs.AI
- Vikas Reddy, Sumanth Challaram. Don’t Ask the LLM to Track Freshness: A Deterministic Recipe for Memory Conflict Resolution. arXiv:2606.01435 — 분야: cs.AI, cs.CL, cs.IR
- Abel Yagubyan. The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation. arXiv:2606.13685 — 분야: cs.CL, cs.AI
- Hongtao Lyu 외. CoAgent: Concurrency Control for Multi-Agent Systems. arXiv:2606.15376 — 분야: cs.DC, cs.AI, cs.MA
- Yanki Margalit 외. Governed Shared Memory for Multi-Agent LLM Systems. arXiv:2606.24535 — 분야: cs.AI
[2026-07-08] 메모리에 무엇이 남았나 — 다운스트림 성공이 아니라 복원 가능성으로 재는 MemProbe
- 중심: Enze Ma 외. MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery. arXiv:2606.24595 — 분야: cs.CL
- Omer Hofman 외. MAPS: A Multilingual Benchmark for Agent Performance and Security. arXiv:2505.15935 — 분야: cs.DB, cs.CL, cs.CR
- Preethi Seshadri 외. Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations. arXiv:2601.17087 — 분야: cs.HC, cs.AI, cs.CY, cs.LG
- Zexue He 외. MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks. arXiv:2602.16313 — 분야: cs.CL
- Chingkwun Lam 외. Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework. arXiv:2603.11768 — 분야: cs.AI
- Shuochen Liu 외. PERMA: Benchmarking Personalized Memory Agents via Event-Driven Preference and Realistic Task Environments. arXiv:2603.23231 — 분야: cs.AI
- Hanxiang Chao 외. STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?. arXiv:2605.06527 — 분야: cs.CL
- Junwei Liao 외. MemQ: Integrating Q-Learning into Self-Evolving Memory Agents over Provenance DAGs. arXiv:2605.08374 — 분야: cs.AI
- Abdelghny Orogat, Essam Mansour. Is Agent Memory a Database? Rethinking Data Foundations for Long-Term AI Agent Memory. arXiv:2605.26252 — 분야: cs.AI, cs.DB
- Zhikai Chen 외. Exploring Cross-Scenario Generality of Agentic Memory Systems: Diagnostics and a Strong Baseline. arXiv:2606.04315 — 분야: cs.AI
- Ziming Wang. TOKI: A Bitemporal Operator Algebra for Contradiction Resolution in LLM-Agent Persistent Memory. arXiv:2606.06240 — 분야: cs.DB, cs.AI
- Laksh Advani. From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents. arXiv:2606.09863 — 분야: cs.LG
- Guanming Liu 외. StreamMemBench: Streaming Evaluation of Agent Memory for Future-Oriented Assistance. arXiv:2606.14571 — 분야: cs.AI
[2026-07-08] 연구 로그 2 — 저울이 저울과 안 맞을 때: judge 이전 실패의 기록
- 중심: Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
[2026-07-07] 메모리는 데이터베이스인가 — 정합성을 궤적의 속성으로 옮기는 GEM의 재설계
- 중심: Abdelghny Orogat, Essam Mansour. Is Agent Memory a Database? Rethinking Data Foundations for Long-Term AI Agent Memory. arXiv:2605.26252 — 분야: cs.AI, cs.DB
- Junwei Liao 외. MemQ: Integrating Q-Learning into Self-Evolving Memory Agents over Provenance DAGs. arXiv:2605.08374 — 분야: cs.AI
- Vikas Reddy, Sumanth Challaram. Don’t Ask the LLM to Track Freshness: A Deterministic Recipe for Memory Conflict Resolution. arXiv:2606.01435 — 분야: cs.AI, cs.CL, cs.IR
- Yaoqi Chen 외. Beyond Semantic Organization: Memory as Execution State Management for Long-Horizon Agents. arXiv:2606.06090 — 분야: cs.AI
- Ziming Wang. TOKI: A Bitemporal Operator Algebra for Contradiction Resolution in LLM-Agent Persistent Memory. arXiv:2606.06240 — 분야: cs.DB, cs.AI
- Yanki Margalit 외. Governed Shared Memory for Multi-Agent LLM Systems. arXiv:2606.24535 — 분야: cs.AI
- Enze Ma 외. MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery. arXiv:2606.24595 — 분야: cs.CL
- Zitong Shi 외. A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory. arXiv:2607.01935 — 분야: cs.AI
[2026-07-06] 학습된 정책은 어디까지 옮겨 다니나 — Memory-R1의 152개 QA쌍과 보상 설계의 힘
- 중심: Sikuan Yan 외. Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning. arXiv:2508.19828 — 분야: cs.CL, cs.MA
- Chuxuan Hu 외. Breaking Barriers: Do Reinforcement Post Training Gains Transfer To Unseen Domains?. arXiv:2506.19733 — 분야: cs.CL
- Daivik Patel, Shrenik Patel. ENGRAM: Effective, Lightweight Memory Orchestration for Conversational Agents. arXiv:2511.12960 — 분야: cs.MA
- Yi Yu 외. Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents. arXiv:2601.01885 — 분야: cs.CL
- Yanwei Yue 외. Mem-T: Densifying Rewards for Long-Horizon Memory Agents. arXiv:2601.23014 — 분야: cs.LG, cs.CL
- Kunvar Thaman. Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use. arXiv:2605.02964 — 분야: cs.LG, cs.AI
- Sikuan Yan 외. Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents. arXiv:2605.21768 — 분야: cs.LG, cs.MA
- Xinjie He 외. What Training Data Teaches RL Memory Agents: An Empirical Study of Curriculum Effects in Memory-Augmented QA. arXiv:2605.23067 — 분야: cs.CL
- Abdelghny Orogat, Essam Mansour. Is Agent Memory a Database? Rethinking Data Foundations for Long-Term AI Agent Memory. arXiv:2605.26252 — 분야: cs.AI, cs.DB
- Adril Putra Merin 외. Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations. arXiv:2606.00832 — 분야: cs.CL
- Zhikai Chen 외. Exploring Cross-Scenario Generality of Agentic Memory Systems: Diagnostics and a Strong Baseline. arXiv:2606.04315 — 분야: cs.AI
- Vedant Patel. Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents. arXiv:2606.27472 — 분야: cs.CL, cs.AI, cs.LG
[2026-07-05] 메모리를 워크로드에 맞춘다는 것 — 에이전트 네이티브 메모리 시스템의 해부와 정렬의 문제
- 중심: Wei Zhou 외. Are We Ready For An Agent-Native Memory System?. arXiv:2606.24775 — 분야: cs.CL, cs.DB, cs.IR
- Sikuan Yan 외. Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning. arXiv:2508.19828 — 분야: cs.CL, cs.MA
- Saad Alqithami. Forgetful but Faithful: A Cognitive Memory Architecture and Benchmark for Privacy-Aware Generative Agents. arXiv:2512.12856 — 분야: cs.AI, cs.LG
- Qizhi Wang. Democratizing GraphRAG: Linear, CPU-Only Graph Retrieval for Multi-Hop QA. arXiv:2602.23372 — 분야: cs.IR, cs.AI, cs.CL
- Han Chen 외. MemForest: An Efficient Agent Memory System with Hierarchical Temporal Indexing. arXiv:2605.23986 — 분야: cs.DB, cs.AI, cs.MA
- Abdelghny Orogat, Essam Mansour. Is Agent Memory a Database? Rethinking Data Foundations for Long-Term AI Agent Memory. arXiv:2605.26252 — 분야: cs.AI, cs.DB
- Adril Putra Merin 외. Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations. arXiv:2606.00832 — 분야: cs.CL
- Yasmine Omri 외. Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads. arXiv:2606.06448 — 분야: cs.AI
[2026-07-04] 메모리를 스킬로 배우다 — AutoMem과 메타기억, 그리고 통제된 실험실의 경계
- 중심: Shengguang Wu 외. AutoMem: Automated Learning of Memory as a Cognitive Skill. arXiv:2607.01224 — 분야: cs.AI, cs.CL, cs.MA
- Sikuan Yan 외. Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning. arXiv:2508.19828 — 분야: cs.CL, cs.MA
- Shomik Jain 외. Interaction Context Often Increases Sycophancy in LLMs. arXiv:2509.12517 — 분야: cs.HC
- Jonggeun Lee 외. Don’t Adapt Small Language Models for Tools; Adapt Tool Schemas to the Models. arXiv:2510.07248 — 분야: cs.CL
- Haozhen Zhang 외. MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents. arXiv:2602.02474 — 분야: cs.CL, cs.AI, cs.LG
- Zhaoxin Feng 외. Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy. arXiv:2603.16643 — 분야: cs.CL
- Md Nayem Uddin 외. From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents. arXiv:2604.20006 — 분야: cs.CL
- Ziyan Liu 외. Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents. arXiv:2605.30159 — 분야: cs.AI
- Adril Putra Merin 외. Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations. arXiv:2606.00832 — 분야: cs.CL
- Wei Zhou 외. Are We Ready For An Agent-Native Memory System?. arXiv:2606.24775 — 분야: cs.CL, cs.DB, cs.IR
[2026-07-03] 아첨이라 부른 것들을 세어 보니 — 파편화된 구인의 분류표와 전문가의 불일치
- 중심: Meryl Ye 외. What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct. arXiv:2605.21778 — 분야: cs.AI
- Myra Cheng 외. ELEPHANT: Measuring and understanding social sycophancy in LLMs. arXiv:2505.13995 — 분야: cs.CL, cs.AI, cs.CY
- Daniel Vennemeyer 외. Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs. arXiv:2509.21305 — 분야: cs.CL
- Itai Shapira 외. How RLHF Amplifies Sycophancy. arXiv:2602.01002 — 분야: cs.AI
[2026-07-03] 연구 로그 1 — 측정기부터 검증합니다: MAST 재측정 파일럿 개시
- 중심: Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
[2026-07-02] 체면을 재는 저울 — Goffman의 face 위에서 사회적 아첨을 네 축으로
- 중심: Myra Cheng 외. ELEPHANT: Measuring and understanding social sycophancy in LLMs. arXiv:2505.13995 — 분야: cs.CL, cs.AI, cs.CY
- Aaron Fanous 외. SycEval: Evaluating LLM Sycophancy. arXiv:2502.08177 — 분야: cs.AI
- Shomik Jain 외. Interaction Context Often Increases Sycophancy in LLMs. arXiv:2509.12517 — 분야: cs.HC
- Myra Cheng 외. Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence. arXiv:2510.01395 — 분야: cs.CY, cs.AI
- Rifo Genadi 외. Sycophancy Hides Linearly in the Attention Heads. arXiv:2601.16644 — 분야: cs.CL, cs.AI
- Zhaoxin Feng 외. Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy. arXiv:2603.16643 — 분야: cs.CL
- Meryl Ye 외. What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct. arXiv:2605.21778 — 분야: cs.AI
[2026-07-01] 아첨을 다섯 항으로 가른다 — 압박 항복과 증거 외면의 분해 보상
- 중심: Muhammad Ahmed Mohsin 외. Pressure, What Pressure? Sycophancy Disentanglement in Language Models via Reward Decomposition. arXiv:2604.05279 — 분야: cs.AI
- Myra Cheng 외. ELEPHANT: Measuring and understanding social sycophancy in LLMs. arXiv:2505.13995 — 분야: cs.CL, cs.AI, cs.CY
- Joy Bhalla, Kristina Gligorić. SWAY: A Counterfactual Computational Linguistic Approach to Measuring and Mitigating Sycophancy. arXiv:2604.02423 — 분야: cs.CL, cs.CY
- William Parris. Semantic Reward Collapse and the Preservation of Epistemic Integrity in Adaptive AI Systems. arXiv:2605.12406 — 분야: cs.AI
- Boyu Xiao 외. When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure. arXiv:2605.23932 — 분야: cs.AI, cs.CL, cs.CY, cs.LG
[2026-06-30] DPO는 언제 RLHF가 아닌가 — 조건부 등가성의 붕괴와 최소 수정
- 중심: Zhiqin Yang 외. Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment. arXiv:2605.20834 — 분야: cs.AI, cs.LG
- Jiancong Xiao 외. On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization. arXiv:2405.16455 — 분야: stat.ML, cs.LG, stat.ME
- Masanari Oi 외. Autoregressive Direct Preference Optimization. arXiv:2602.09533 — 분야: cs.AI
- Suqin Yuan 외. Mitigating Mismatch within Reference-based Preference Optimization. arXiv:2602.11902 — 분야: cs.LG, cs.AI
- Xiaoyi Li. Do Post-Training Algorithms Actually Differ? A Controlled Study Across Model Scales Uncovers Scale-Dependent Ranking Inversions. arXiv:2603.19335 — 분야: cs.LG, cs.AI
- Muhammad Ahmed Mohsin 외. Pressure, What Pressure? Sycophancy Disentanglement in Language Models via Reward Decomposition. arXiv:2604.05279 — 분야: cs.AI
[2026-06-29] 훈련이 아첨을 키운다 — RLHF 공분산 증폭과 최소 교정
- 중심: Itai Shapira 외. How RLHF Amplifies Sycophancy. arXiv:2602.01002 — 분야: cs.AI
- Myra Cheng 외. ELEPHANT: Measuring and understanding social sycophancy in LLMs. arXiv:2505.13995 — 분야: cs.CL, cs.AI, cs.CY
- Daniel Fein 외. One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models. arXiv:2603.03291 — 분야: cs.CL, cs.AI
- Muhammad Ahmed Mohsin 외. Pressure, What Pressure? Sycophancy Disentanglement in Language Models via Reward Decomposition. arXiv:2604.05279 — 분야: cs.AI
- Zhiqin Yang 외. Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment. arXiv:2605.20834 — 분야: cs.AI, cs.LG
- Meryl Ye 외. What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct. arXiv:2605.21778 — 분야: cs.AI
[2026-06-28] 아첨은 하나가 아니다 — SyA·GA·SyPR의 인과적 분리
- 중심: Daniel Vennemeyer 외. Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs. arXiv:2509.21305 — 분야: cs.CL
- Itai Shapira 외. How RLHF Amplifies Sycophancy. arXiv:2602.01002 — 분야: cs.AI
- Cansu Koyuturk 외. The Hidden Cost of Contextual Sycophancy: an AI Literacy Intervention in Human-AI Collaboration. arXiv:2605.18372 — 분야: cs.HC, cs.AI, cs.CY, cs.ET
- Meryl Ye 외. What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct. arXiv:2605.21778 — 분야: cs.AI
[2026-06-27] 아첨이 친절을 줄인다 — 사회적 아첨은 관계 수리 의지를 깎고 의존을 키운다
- 중심: Myra Cheng 외. Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence. arXiv:2510.01395 — 분야: cs.CY, cs.AI
- Daniel Vennemeyer 외. Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs. arXiv:2509.21305 — 분야: cs.CL
- Cansu Koyuturk 외. The Hidden Cost of Contextual Sycophancy: an AI Literacy Intervention in Human-AI Collaboration. arXiv:2605.18372 — 분야: cs.HC, cs.AI, cs.CY, cs.ET
- Meryl Ye 외. What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct. arXiv:2605.21778 — 분야: cs.AI
[2026-06-26] 합리적이어도 빠진다 — 아첨하는 챗봇은 이상적 베이지안조차 망상으로 끌고 간다
- 중심: Kartik Chandra 외. Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians. arXiv:2602.19141 — 분야: cs.AI, cs.CY, cs.HC
- Sebastian Dohnány 외. Technological folie à deux: Feedback Loops Between AI Chatbots and Mental Illness. arXiv:2507.19218 — 분야: cs.HC, cs.AI, q-bio.NC
- Myra Cheng 외. Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence. arXiv:2510.01395 — 분야: cs.CY, cs.AI
- Meryl Ye 외. What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct. arXiv:2605.21778 — 분야: cs.AI
[2026-06-25] 덮어쓰이는 진실 — 아첨은 저장된 편향이 아니라 후기 레이어의 생성물이다
- 중심: Keyu Wang 외. When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models. arXiv:2508.02087 — 분야: cs.CL
- Daniel Vennemeyer 외. Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs. arXiv:2509.21305 — 분야: cs.CL
- Rifo Genadi 외. Sycophancy Hides Linearly in the Attention Heads. arXiv:2601.16644 — 분야: cs.CL, cs.AI
- Claire O’Brien 외. A Few Bad Neurons: Isolating and Surgically Correcting Sycophancy. arXiv:2601.18939 — 분야: cs.LG
- Itai Shapira 외. How RLHF Amplifies Sycophancy. arXiv:2602.01002 — 분야: cs.AI
- Kartik Chandra 외. Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians. arXiv:2602.19141 — 분야: cs.AI, cs.CY, cs.HC
- Petter Törnberg, Michelle Schimmel. Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor. arXiv:2604.27633 — 분야: cs.AI
- Adarsh Kumarappan, Ananya Mujoo. Not Just RLHF: Why Alignment Alone Won’t Fix Multi-Agent Sycophancy. arXiv:2605.12991 — 분야: cs.LG, cs.AI
[2026-06-24] 전제로 굳은 의심 — 편향을 판단하는 회로가 기울 때
- 중심: Ramaravind Kommiya Mothilal 외. Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement. arXiv:2606.17506 — 분야: cs.CL
- Xuyang Wu 외. Does Reasoning Introduce Bias? A Study of Social Bias Evaluation and Mitigation in LLM Reasoning. arXiv:2502.15361 — 분야: cs.CL, cs.AI
- Tiansheng Huang 외. Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable. arXiv:2503.00555 — 분야: cs.CR, cs.AI, cs.LG
- Qingquan Li 외. Evaluating Scoring Bias in LLM-as-a-Judge. arXiv:2506.22316 — 분야: cs.CL
- Srikant Panda 외. DAIQ: Auditing Demographic Attribute Inference from Question in LLMs. arXiv:2508.15830 — 분야: cs.CL, cs.AI
- Rom Himelstein 외. Silenced Biases: The Dark Side LLMs Learned to Refuse. arXiv:2511.03369 — 분야: cs.CL, stat.ML
- Xiaolin Zhou 외. Fairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-Judge. arXiv:2601.13649 — 분야: cs.CL, cs.AI
- Zixiao Zhao 외. Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering. arXiv:2604.16790 — 분야: cs.SE, cs.AI
- Edie Pearman 외. Mechanics of Bias and Reasoning: Interpreting the Impact of Chain-of-Thought Prompting on Gender Bias in LLMs. arXiv:2605.20410 — 분야: cs.CL, cs.AI
[2026-06-23] 중립의 환상 — 편향이 없어 보이는 것과 평가할 줄 모르는 것
- 중심: Kevin T Webster. Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening. arXiv:2507.11548 — 분야: cs.CY, cs.AI, cs.CL
- Eitan Anzenberg 외. Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions. arXiv:2507.02087 — 분야: cs.LG, cs.CL, cs.CY
- Honglin Mu 외. AI Security Beyond Core Domains: Resume Screening as a Case Study of Adversarial Vulnerabilities in Specialized LLM Applications. arXiv:2512.20164 — 분야: cs.CL, cs.AI
- José Pombal 외. Self-Preference Bias in Rubric-Based Evaluation of Large Language Models. arXiv:2604.06996 — 분야: cs.CL, cs.AI
- Ramaravind Kommiya Mothilal 외. Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement. arXiv:2606.17506 — 분야: cs.CL
[2026-06-22] 거울을 깨는 한 방향 — 유해 자기선호만 또렷한 선, 정당 편애는 흩어진 안개
- 중심: Dani Roytburg 외. Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators. arXiv:2509.03647 — 분야: cs.CL, cs.AI, cs.LG
- Daniel Tan 외. Analyzing the Generalization and Reliability of Steering Vectors. arXiv:2407.12404 — 분야: cs.LG
- Koki Wataoka 외. Self-Preference Bias in LLM-as-a-Judge. arXiv:2410.21819 — 분야: cs.CL
- Vincent Siu 외. SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs. arXiv:2509.13450 — 분야: cs.AI, cs.CL, cs.LG
- Steven A. Lehr 외. Extreme Self-Preference in Language Models. arXiv:2509.26464 — 분야: cs.AI, cs.CL, cs.LG
- Tim Tian Hua 외. Steering Evaluation-Aware Language Models to Act Like They Are Deployed. arXiv:2510.20487 — 분야: cs.CL, cs.AI
- José Pombal 외. Self-Preference Bias in Rubric-Based Evaluation of Large Language Models. arXiv:2604.06996 — 분야: cs.CL, cs.AI
- Jinming Yang 외. Quantifying and Mitigating Self-Preference Bias of LLM Judges. arXiv:2604.22891 — 분야: cs.LG, cs.AI, cs.CL
[2026-06-21] 이유 있는 편애와 이유 없는 고집 — 강한 심판이 틀릴 때 가장 깊어지는 맹점
- 중심: Wei-Lin Chen 외. Do LLM Evaluators Prefer Themselves for a Reason?. arXiv:2504.03846 — 분야: cs.CL
- Arjun Panickssery 외. LLM Evaluators Recognize and Favor Their Own Generations. arXiv:2404.13076 — 분야: cs.CL, cs.AI
- Koki Wataoka 외. Self-Preference Bias in LLM-as-a-Judge. arXiv:2410.21819 — 분야: cs.CL
- Dani Roytburg 외. Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators. arXiv:2509.03647 — 분야: cs.CL, cs.AI, cs.LG
- José Pombal 외. Self-Preference Bias in Rubric-Based Evaluation of Large Language Models. arXiv:2604.06996 — 분야: cs.CL, cs.AI
- William Guey, Pierrick Bougault. Self-Preference Is Weak or Absent in Verifiable Instruction-Following Revision: A Four-Model Test Under Genuine Authorship. arXiv:2606.20093 — 분야: cs.CL
[2026-06-20] 내 이력서를 내가 뽑는다 — LLM 자기선호가 채용 파이프라인을 잠그는 법
- 중심: Jiannan Xu 외. AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights. arXiv:2509.00462 — 분야: cs.CY
- Wei-Lin Chen 외. Do LLM Evaluators Prefer Themselves for a Reason?. arXiv:2504.03846 — 분야: cs.CL
- Kevin T Webster. Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening. arXiv:2507.11548 — 분야: cs.CY, cs.AI, cs.CL
- Dani Roytburg 외. Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators. arXiv:2509.03647 — 분야: cs.CL, cs.AI, cs.LG
- Jinming Yang 외. Quantifying and Mitigating Self-Preference Bias of LLM Judges. arXiv:2604.22891 — 분야: cs.LG, cs.AI, cs.CL
[2026-06-19] 닮아가는 오답들 — 더 똑똑한 모델일수록 같은 자리에서 함께 틀린다
- 중심: Elliot Kim 외. Correlated Errors in Large Language Models. arXiv:2506.07962 — 분야: cs.CL, cs.AI, cs.CY, stat.ML
- Shashwat Goel 외. Great Models Think Alike and this Undermines AI Oversight. arXiv:2502.04313 — 분야: cs.LG, cs.AI, cs.CL
- Jiannan Xu 외. AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights. arXiv:2509.00462 — 분야: cs.CY
- Dustin Wright 외. Epistemic Diversity and Knowledge Collapse in Large Language Models. arXiv:2510.04226 — 분야: cs.CL, cs.AI, cs.CY, cs.IR, cs.LG
- Yingxuan Yang 외. Understanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity. arXiv:2602.03794 — 분야: cs.AI, cs.LG
- Geunbin Yu. AdaptOrch: Task-Adaptive Multi-Agent Orchestration in the Era of LLM Performance Convergence. arXiv:2602.16873 — 분야: cs.MA, cs.AI
- Nathanael Jo 외. The Subjectivity of Monoculture. arXiv:2602.24086 — 분야: cs.CY, cs.LG
[2026-06-18] 공동 실패를 어렵게 짓는다 — Council Mode는 이질 합의를 구조로 설계한다
- 중심: Shuai Wu 외. Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias. arXiv:2604.02923 — 분야: cs.CL, cs.AI
- Wenzhe Li 외. Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?. arXiv:2502.00674 — 분야: cs.CL, cs.LG
- Elliot Kim 외. Correlated Errors in Large Language Models. arXiv:2506.07962 — 분야: cs.CL, cs.AI, cs.CY, stat.ML
- Wenting Zhao 외. The Majority is not always right: RL training for solution aggregation. arXiv:2509.06870 — 분야: cs.CL
- Antonio Sabbatella. MALBO: Optimizing LLM-Based Multi-Agent Teams via Multi-Objective Bayesian Optimization. arXiv:2511.11788 — 분야: cs.MA, cs.AI
- Yubin Kim 외. Towards a Science of Scaling Agent Systems. arXiv:2512.08296 — 분야: cs.AI
- Wei Yang 외. Auditing Multi-Agent LLM Reasoning Trees Outperforms Majority Vote and LLM-as-Judge. arXiv:2602.09341 — 분야: cs.AI
- Zhuo Li 외. MARCH: Multi-Agent Reinforced Self-Check for LLM Hallucination. arXiv:2603.24579 — 분야: cs.CL
- Michał Wawer, Jarosław A. Chudziak. Consensus is Strategically Insufficient: Reasoning-Trace Disagreement as a Knowledge-Representation Signal. arXiv:2606.04223 — 분야: cs.AI
[2026-06-17] 모델은 자기가 틀린 걸 알까 — 숨겨진 상태는 진실이 아니라 회상을 비춘다
- 중심: Chi Seng Cheang 외. Do LLMs Really Know What They Don’t Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness. arXiv:2510.09033 — 분야: cs.CL
- Hadas Orgad 외. LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations. arXiv:2410.02707 — 분야: cs.CL, cs.AI
- Keyu Wang 외. When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models. arXiv:2508.02087 — 분야: cs.CL
- Shaowen Wang 외. When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs. arXiv:2511.07318 — 분야: cs.CL, cs.AI, cs.LG
- Khizar Hussain, Murat Kantarcioglu. PARALLAX: Separating Genuine Hallucination Detection from Benchmark Construction Artifacts. arXiv:2605.17028 — 분야: cs.CL, cs.AI
[2026-06-16] 답이 맞아도 이유는 달랐다 — 합의가 가린 것을 CARA가 재는 법
- 중심: Xiaoyang Wang, Christopher C. Yang. The Consistency Illusion: How Multi-Agent Debate Hides Reasoning Misalignment. arXiv:2606.08457 — 분야: cs.MA
- Andrea Wynn 외. Talk Isn’t Always Cheap: Understanding Failure Modes in Multi-Agent Debate. arXiv:2509.05396 — 분야: cs.CL, cs.AI, cs.MA
- Blaž Bertalanič, Carolina Fortuna. The Cost of Consensus: Isolated Self-Correction Prevails Over Unguided Homogeneous Multi-Agent Debate. arXiv:2605.00914 — 분야: cs.MA, cs.AI
- Michał Wawer, Jarosław A. Chudziak. Consensus is Strategically Insufficient: Reasoning-Trace Disagreement as a Knowledge-Representation Signal. arXiv:2606.04223 — 분야: cs.AI
[2026-06-15] 잠입자를 찾아내면 합의가 깨끗해질까 — MUG는 환각하는 에이전트를 반사실로 색출한다
- 중심: Dayong Liang 외. Multi-agent Undercover Gaming: Hallucination Removal via Counterfactual Test for Multimodal Reasoning. arXiv:2511.11182 — 분야: cs.AI, cs.CL, cs.MA, cs.MM
- Yijun Feng. Counterfactual Probing for Hallucination Detection and Mitigation in Large Language Models. arXiv:2508.01862 — 분야: cs.CL, cs.AI
- Xuannan Liu 외. AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents. arXiv:2601.06818 — 분야: cs.CL
- Bang Liu 외. Phase Transition for Budgeted Multi-Agent Synergy. arXiv:2601.17311 — 분야: cs.AI
- Shuai Wu 외. Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias. arXiv:2604.02923 — 분야: cs.CL, cs.AI
- Xiaoyang Wang, Christopher C. Yang. The Consistency Illusion: How Multi-Agent Debate Hides Reasoning Misalignment. arXiv:2606.08457 — 분야: cs.MA
[2026-06-14] 빈 우물이 아니라 잘못 잡은 삽이었다면 — MechELK는 표면 아래 잠긴 지식을 인과로 길어 올린다
- 중심: Ji-jun Park 외. MechELK: A Mechanistic Interpretability Framework for Eliciting Latent Knowledge in Large Language Models. arXiv:2605.28825 — 분야: cs.CL
- Stefan F. Schouten 외. Truth-value judgment in language models: ‘truth directions’ are context sensitive. arXiv:2404.18865 — 분야: cs.CL
- Joseph Miller 외. Transformer Circuit Faithfulness Metrics are not Robust. arXiv:2407.08734 — 분야: cs.LG, cs.AI, cs.CL
- Daniel Tan 외. Analyzing the Generalization and Reliability of Steering Vectors. arXiv:2407.12404 — 분야: cs.LG
- Hadas Orgad 외. LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations. arXiv:2410.02707 — 분야: cs.CL, cs.AI
- Yu Zhao 외. Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering. arXiv:2410.15999 — 분야: cs.CL
- Jingyi Cui 외. On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted Remedy. arXiv:2506.15963 — 분야: cs.LG
- Chi Seng Cheang 외. Do LLMs Really Know What They Don’t Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness. arXiv:2510.09033 — 분야: cs.CL
- Anton Korznikov 외. Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?. arXiv:2602.14111 — 분야: cs.LG
- David Chanin. Are Sparse Autoencoder Benchmarks Reliable?. arXiv:2605.18229 — 분야: cs.LG, cs.AI
- Xinpeng Wang 외. Automatic Layer Selection for Hallucination Detection. arXiv:2605.26366 — 분야: cs.AI, cs.LG
[2026-06-13] 직관이 가리킨 곳을 파보니 빈 우물이었다 — 환각과 지식 충돌은 내부 표현에서 만나지 않는다
- 중심: Lucrezia Laraspata 외. Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models. arXiv:2606.08705 — 분야: cs.CL
- Muru Zhang 외. How Language Model Hallucinations Can Snowball. arXiv:2305.13534 — 분야: cs.CL
- Yufei Tao 외. When Context Leads but Parametric Memory Follows in Large Language Models. arXiv:2409.08435 — 분야: cs.CL, cs.AI
- Yu Zhao 외. Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering. arXiv:2410.15999 — 분야: cs.CL
- Yu Zhao 외. Analysing the Residual Stream of Language Models Under Knowledge Conflicts. arXiv:2410.16090 — 분야: cs.CL
- Zuzanna Dubanowska 외. Representation-based Broad Hallucination Detectors Fail to Generalize Out of Distribution. arXiv:2509.19372 — 분야: cs.LG, cs.AI
- Adrian Robert Minut 외. Spilled Energy in Large Language Models. arXiv:2602.18671 — 분야: cs.AI, cs.CL
- Shanshan Lin 외. Constrained Paraphrase Consistency for LLM Hallucination Detection. arXiv:2606.08158 — 분야: cs.CL, cs.AI
[2026-06-12] 환각은 출력에 머물지 않고 연쇄를 따라 흐른다 — Hallucination Cascade가 본 전파의 동역학
- 중심: Saeid Jamshidi 외. Hallucination Cascade: Analyzing Error Propagation in Multi-Agent LLM Systems. arXiv:2606.07937 — 분야: cs.CR
- Dayong Liang 외. Multi-agent Undercover Gaming: Hallucination Removal via Counterfactual Test for Multimodal Reasoning. arXiv:2511.11182 — 분야: cs.AI, cs.CL, cs.MA, cs.MM
- Xuannan Liu 외. AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents. arXiv:2601.06818 — 분야: cs.CL
- Naen Xu 외. When Agents “Misremember” Collectively: Exploring the Mandela Effect in LLM-based Multi-Agent Systems. arXiv:2602.00428 — 분야: cs.CL, cs.AI, cs.CR
- Yawen Wang 외. From Flat Logs to Causal Graphs: Hierarchical Failure Attribution for LLM-based Multi-Agent Systems. arXiv:2602.23701 — 분야: cs.AI, cs.SE
- Yizhe Xie 외. From Spark to Fire: Modeling and Mitigating Error Cascades in LLM-Based Multi-Agent Collaboration. arXiv:2603.04474 — 분야: cs.MA, cs.AI
- Shuai Wu 외. Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias. arXiv:2604.02923 — 분야: cs.CL, cs.AI
- Xiaoyang Wang, Christopher C. Yang. The Consistency Illusion: How Multi-Agent Debate Hides Reasoning Misalignment. arXiv:2606.08457 — 분야: cs.MA
[2026-06-11] 장부를 쥔 손이 장부를 고쳐 쓸 때 — Self-Harness가 에이전트에게 자기 하니스를 맡기는 법
- 중심: Hangfan Zhang 외. Self-Harness: Harnesses That Improve Themselves. arXiv:2606.09498 — 분야: cs.CL
- Yoonho Lee 외. Meta-Harness: End-to-End Optimization of Model Harnesses. arXiv:2603.28052 — 분야: cs.AI
- Jiahang Lin 외. Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses. arXiv:2604.25850 — 분야: cs.CL, cs.SE
- Yong-eun Cho. It’s Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers. arXiv:2605.26731 — 분야: cs.AI, cs.CL
- Prannay Hebbar 외. SIA: Self Improving AI with Harness & Weight Updates. arXiv:2605.27276 — 분야: cs.AI, cs.CL
- Minhua Lin 외. Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents. arXiv:2605.30621 — 분야: cs.AI
- Wenbo Pan 외. Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference. arXiv:2606.05922 — 분야: cs.AI, cs.CL, cs.LG
- Mengzhuo Chen 외. From Failed Trajectories to Reliable LLM Agents: Diagnosing and Repairing Harness Flaws. arXiv:2606.06324 — 분야: cs.SE, cs.MA
[2026-06-10] 이름 붙인 자리에 붕대를 두르는 일 — FAMA가 실패에서 최소한의 손길만 골라내는 법
- 중심: Amir Saeidi 외. FAMA: Failure-Aware Meta-Agentic Framework for Open-Source LLMs in Interactive Tool Use Environments. arXiv:2604.25135 — 분야: cs.CL
- Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
- Venkatesh Mishra 외. How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $τ$-bench. arXiv:2508.20931 — 분야: cs.CL
- Sri Vatsa Vuddanti 외. PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases. arXiv:2509.25238 — 분야: cs.LG, cs.AI
- JV Roig. How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations. arXiv:2512.07497 — 분야: cs.AI, cs.SE
- Geunbin Yu. AdaptOrch: Task-Adaptive Multi-Agent Orchestration in the Era of LLM Performance Convergence. arXiv:2602.16873 — 분야: cs.MA, cs.AI
- Dat Tran, Douwe Kiela. Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets. arXiv:2604.02460 — 분야: cs.CL, cs.MA
- Mengzhuo Chen 외. Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems. arXiv:2604.22708 — 분야: cs.MA
[2026-06-09] 무너지는 자리에 이름을 붙이는 일 — MAST가 다중 에이전트 시스템의 실패를 해부하는 법
- 중심: Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
- Shanshan Han 외. LLM Multi-Agent Systems: Challenges and Open Problems. arXiv:2402.03578 — 분야: cs.MA, cs.AI
- Lewis Hammond 외. Multi-Agent Risks from Advanced AI. arXiv:2502.14143 — 분야: cs.MA, cs.AI, cs.CY, cs.ET, cs.LG
- Shaokun Zhang 외. Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems. arXiv:2505.00212 — 분야: cs.MA, cs.CL
- Yuxuan Li 외. Systematic Failures in Collective Reasoning under Distributed Information in Multi-Agent LLMs. arXiv:2505.11556 — 분야: cs.CL, cs.AI, cs.MA
- Aaron Xuxiang Tian 외. Beyond the Strongest LLM: Multi-Turn Multi-Agent Orchestration vs. Single LLMs on Benchmarks. arXiv:2509.23537 — 분야: cs.AI
- Khush Patel 외. The Six Sigma Agent: Achieving Enterprise-Grade Reliability in LLM Systems Through Consensus-Driven Decomposed Execution. arXiv:2601.22290 — 분야: cs.AI
[2026-06-08] 에이전트가 에이전트를 짜는 날 — MAC가 벤치마크에 없던 질문을 던지다
- 중심: Xinyu Lu 외. The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?. arXiv:2606.04455 — 분야: cs.AI, cs.CL
- Shengran Hu 외. Automated Design of Agentic Systems. arXiv:2408.08435 — 분야: cs.AI
- Govind Pimpale 외. Forecasting Frontier Language Model Agent Capabilities. arXiv:2502.15850 — 분야: cs.CL, cs.AI
- Mert Cemri 외. Why Do Multi-Agent LLM Systems Fail?. arXiv:2503.13657 — 분야: cs.AI
- Maxime Robeyns 외. A Self-Improving Coding Agent. arXiv:2504.15228 — 분야: cs.AI
- Hongjin Qian, Zheng Liu. MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning. arXiv:2508.00271 — 분야: cs.AI, cs.CL, cs.IR
- Shuai Shao 외. Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents. arXiv:2509.26354 — 분야: cs.AI, cs.CL, cs.LG
- Monte MacDiarmid 외. Natural Emergent Misalignment from Reward Hacking in Production RL. arXiv:2511.18397 — 분야: cs.AI, cs.SE
- Darshan Deshpande 외. Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis. arXiv:2601.20103 — 분야: cs.SE, cs.AI, cs.LG
- Ben Rank 외. PostTrainBench: Can LLM Agents Automate LLM Post-Training?. arXiv:2603.08640 — 분야: cs.SE, cs.AI, cs.LG
- Kunvar Thaman. Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use. arXiv:2605.02964 — 분야: cs.LG, cs.AI
[2026-06-07] 루브릭이 공유 인터페이스가 될 때 — RubricEM이 정책·판사·기억을 하나로 묶는 방식
- 중심: Gaotang Li 외. RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards. arXiv:2605.10899 — 분야: cs.CL, cs.LG
- Quan Wei 외. Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Reward Design. arXiv:2505.11821 — 분야: cs.LG
- Rulin Shao 외. DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research. arXiv:2511.19399 — 분야: cs.CL, cs.AI, cs.LG
- Hui-Ze Tan 외. Hindsight Credit Assignment for Long-Horizon LLM Agents. arXiv:2603.08754 — 분야: cs.LG, cs.AI
- Teng Xiao 외. Meta-Reinforcement Learning with Self-Reflection for Agentic Search. arXiv:2603.11327 — 분야: cs.LG, cs.CL
- Liang Ding. AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning. arXiv:2603.21362 — 분야: cs.AI, cs.CL
- José Pombal 외. Self-Preference Bias in Rubric-Based Evaluation of Large Language Models. arXiv:2604.06996 — 분야: cs.CL, cs.AI
- Hao Han 외. SWE-TRACE: Optimizing Long-Horizon SWE Agents Through Rubric Process Reward Models and Heuristic Test-Time Scaling. arXiv:2604.14820 — 분야: cs.SE
- Hongyi Liu 외. SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution. arXiv:2605.18401 — 분야: cs.CL, cs.AI
- Xuekang Wang 외. Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning. arXiv:2606.04923 — 분야: cs.LG, cs.AI, cs.CL
[2026-06-06] 기준의 탄생을 누가 결정하나 — ARES가 사전훈련 문서에서 루브릭을 길어 올리는 법
- 중심: Xiaoyuan Li 외. ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning. arXiv:2605.23454 — 분야: cs.CL
- Pengkai Wang 외. InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training. arXiv:2510.15859 — 분야: cs.CL, cs.AI
- Ran Xu 외. Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training. arXiv:2602.01511 — 분야: cs.CL, cs.LG
- William F. Shen 외. Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks. arXiv:2602.05125 — 분야: cs.LG, cs.AI
- Gaotang Li 외. RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards. arXiv:2605.10899 — 분야: cs.CL, cs.LG
- Anas Mahmoud 외. Reward Hacking in Rubric-Based Reinforcement Learning. arXiv:2605.12474 — 분야: cs.AI
[2026-06-05] 기준을 정책이 들지 않는다, 메모리가 들고 키운다 — ARBOR가 process reward를 살려두는 법
- 중심: Zheng Liu 외. ARBOR: Online Process Rewards via a Reusable Rubric Buffer for Search Agents. arXiv:2606.03239 — 분야: cs.CL
- Jiaxuan Gao 외. On Designing Effective RL Reward at Training Time for LLM Reasoning. arXiv:2410.15115 — 분야: cs.LG, cs.AI, cs.CL
- Chenlu Ye 외. Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training. arXiv:2509.03403 — 분야: cs.LG, cs.AI
- Mingkang Zhu 외. Stratified GRPO: Handling Structural Heterogeneity in Reinforcement Learning of LLM Search Agents. arXiv:2510.06214 — 분야: cs.LG, cs.AI, cs.CL
- Ran Xu 외. Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training. arXiv:2602.01511 — 분야: cs.CL, cs.LG
- Zhi Zhang 외. Train Less, Learn More: Adaptive Efficient Rollout Optimization for Group-Based Reinforcement Learning. arXiv:2602.14338 — 분야: cs.LG, cs.AI
- Xinyu Wang 외. Co-Evolution of Policy and Internal Reward for Language Agents. arXiv:2604.03098 — 분야: cs.LG, cs.AI, cs.CL
- Xiaoyuan Li 외. ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning. arXiv:2605.23454 — 분야: cs.CL
- Nianyi Lin 외. LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards. arXiv:2605.31584 — 분야: cs.CL, cs.AI, cs.LG
[2026-06-04] 정책은 결정만 하라, 장부는 환경이 쥔다 — Harness-1이 검색 상태를 외부화하는 방식
- 중심: Pengcheng Jiang 외. Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses. arXiv:2606.02373 — 분야: cs.AI, cs.CL, cs.IR
- Sikuan Yan 외. Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning. arXiv:2508.19828 — 분야: cs.CL, cs.MA
- Yuxiang Ji 외. Tree Search for LLM Agent Reinforcement Learning. arXiv:2509.21240 — 분야: cs.LG, cs.AI
- Yiding Wang 외. Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents. arXiv:2510.04695 — 분야: cs.AI
- Yibo Zhao 외. Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?. arXiv:2605.27881 — 분야: cs.CL
[2026-06-03] 맞은 답에도 새는 곳이 있다 — TELBench·DRIFT가 궤적에서 오류의 발원지를 짚는 법
- 중심: Jiaming Wang 외. Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories. arXiv:2606.02060 — 분야: cs.AI
- Yindong Wang 외. ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations. arXiv:2509.25868 — 분야: cs.CL
- Youliang Yuan 외. Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards. arXiv:2510.07774 — 분야: cs.CL
- Zhiheng Xi 외. AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress. arXiv:2511.08325 — 분야: cs.CL, cs.IR, cs.LG
- Donald Ye 외. Mechanistic Evidence for Faithfulness Decay in Chain-of-Thought Reasoning. arXiv:2602.11201 — 분야: cs.CL
- Zhisong Qiu 외. Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis. arXiv:2604.24198 — 분야: cs.CL, cs.AI, cs.CE, cs.LG, cs.MA
- Harshada Badave 외. Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows. arXiv:2605.24219 — 분야: cs.AI
[2026-06-02] 검색은 이겼는데 천장은 같다 — PROBE가 프로액티브 에이전트를 세 조각으로 해부하는 방식
- 중심: Gil Pasternak 외. Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents. arXiv:2510.19771 — 분야: cs.AI
- Mudit Verma 외. On the Brittle Foundations of ReAct Prompting for Agentic Large Language Models. arXiv:2405.13966 — 분야: cs.AI, cs.CL
- Taiming Lu 외. Insights into LLM Long-Context Failures: When Transformers Know but Don’t Tell. arXiv:2406.14673 — 분야: cs.CL
- Shaokun Zhang 외. Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems. arXiv:2505.00212 — 분야: cs.MA, cs.CL
- Yuanbo Tang 외. ProAgentBench: Evaluating LLM Agents for Proactive Assistance with Real-World Data. arXiv:2602.04482 — 분야: cs.HC
- Deepak Nathani 외. Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants. arXiv:2604.00842 — 분야: cs.AI, cs.LG, cs.MA
- Mengzhuo Chen 외. Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems. arXiv:2604.22708 — 분야: cs.MA
- Haoming Meng. CUJBench: Benchmarking LLM-Agent on Cross-Modal Failure Diagnosis from Browser to Backend. arXiv:2604.23455 — 분야: cs.SE
[2026-06-01] 출처를 기억하는 그래프 — MemORAI가 대화 메모리에 이력을 새기는 방식
- 중심: Hung Pham Van 외. MemORAI: Memory Organization and Retrieval via Adaptive Graph Intelligence for LLM Conversational Agents. arXiv:2605.01386 — 분야: cs.CL
- Hithesh Sankararaman 외. Provenance: A Light-weight Fact-checker for Retrieval Augmented LLM Generation Output. arXiv:2411.01022 — 분야: cs.CL
- Qingyao Ai 외. MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems. arXiv:2510.17281 — 분야: cs.LG, cs.AI, cs.IR
- Gil Pasternak 외. Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents. arXiv:2510.19771 — 분야: cs.AI
- Daniel Herbst 외. Lost in Serialization: Invariance and Generalization of LLM Graph Reasoners. arXiv:2511.10234 — 분야: cs.LG, cs.AI
- Michael H. Coen. When F1 Fails: Granularity-Aware Evaluation for Dialogue Topic Segmentation. arXiv:2512.17083 — 분야: cs.CL, cs.AI
- Wenyu Mao 외. Bi-Mem: Bidirectional Construction of Hierarchical Memory for Personalized LLMs via Inductive-Reflective Agents. arXiv:2601.06490 — 분야: cs.MA
- Swarna Kamal Paul 외. GAAMA: Graph Augmented Associative Memory for Agents. arXiv:2603.27910 — 분야: cs.AI, cs.IR, cs.MA
- Zhaofen Wu 외. GAM: Hierarchical Graph-based Agentic Memory for LLM Agents. arXiv:2604.12285 — 분야: cs.AI
[2026-05-31] 깨어날 때를 누가 정하는가 — 프로액티브 에이전트의 트리거를 그래프에 돌려주다
- 중심: Xiaoze Liu 외. Do Proactive Agents Really Need an LLM to Decide When to Wake and What to Anchor?. arXiv:2605.30152 — 분야: cs.CL, cs.AI, cs.HC
- Weilin Cong 외. On the Generalization Capability of Temporal Graph Learning Algorithms: Theoretical Insights and a Simpler Method. arXiv:2402.16387 — 분야: cs.LG, cs.AI
- Gil Pasternak 외. Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents. arXiv:2510.19771 — 분야: cs.AI
- Daniel Herbst 외. Lost in Serialization: Invariance and Generalization of LLM Graph Reasoners. arXiv:2511.10234 — 분야: cs.LG, cs.AI
- Yuxuan Fu 외. PRISM: Festina Lente Proactivity – Risk-Sensitive, Uncertainty-Aware Deliberation for Proactive Agents. arXiv:2602.01532 — 분야: cs.AI, cs.HC
- Yuanbo Tang 외. ProAgentBench: Evaluating LLM Agents for Proactive Assistance with Real-World Data. arXiv:2602.04482 — 분야: cs.HC
- Warren Johnson, Charles Lee. Evaluating Small Language Models for Front-Door Routing: A Harmonized Benchmark and Synthetic-Traffic Experiment. arXiv:2604.02367 — 분야: cs.NI, cs.CL
- Hung Pham Van 외. MemORAI: Memory Organization and Retrieval via Adaptive Graph Intelligence for LLM Conversational Agents. arXiv:2605.01386 — 분야: cs.CL
[2026-05-30] 위상은 한 번에 굳지 않는다 — FluxMem이 메모리 그래프를 흐르게 두는 방식
- 중심: Jizhan Fang 외. Rethinking Memory as Continuously Evolving Connectivity. arXiv:2605.28773 — 분야: cs.CL, cs.AI, cs.LG, cs.MA, cs.MM
- Kevin Lin 외. Sleep-time Compute: Beyond Inference Scaling at Test-time. arXiv:2504.13171 — 분야: cs.AI, cs.CL
- Jiaqi Liu 외. SimpleMem: Efficient Lifelong Memory for LLM Agents. arXiv:2601.02553 — 분야: cs.AI
- Chingkwun Lam 외. Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework. arXiv:2603.11768 — 분야: cs.AI
- Can Lv 외. All-Mem: Agentic Lifelong Memory via Dynamic Topology Evolution. arXiv:2603.19595 — 분야: cs.IR, cs.CL
- Hung Pham Van 외. MemORAI: Memory Organization and Retrieval via Adaptive Graph Intelligence for LLM Conversational Agents. arXiv:2605.01386 — 분야: cs.CL
[2026-05-29] 에이전트는 조용히 늙는다 — 배포 후 신뢰성을 라이프스팬으로 측정한다는 것
- 중심: Jianing Zhu 외. Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems. arXiv:2605.26302 — 분야: cs.AI, cs.CL, cs.MA
- Kevin Lin 외. Sleep-time Compute: Beyond Inference Scaling at Test-time. arXiv:2504.13171 — 분야: cs.AI, cs.CL
- Murali Sridharan 외. Detection, Classification and Prevalence of Self-Admitted Aging Debt. arXiv:2504.17428 — 분야: cs.SE, cs.AI, cs.CE, cs.GL
- Shuochen Liu 외. PERMA: Benchmarking Personalized Memory Agents via Event-Driven Preference and Realistic Task Environments. arXiv:2603.23231 — 분야: cs.AI
- Hyunji Lee 외. MINTEval: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems. arXiv:2605.18565 — 분야: cs.CL, cs.AI
[2026-05-28] 기억은 한 번에 저장되지 않는다 — 수면 공고화로 다시 읽는 fast weight 병목
- 중심: Sangyun Lee 외. Language Models Need Sleep. arXiv:2605.26099 — 분야: cs.CL, cs.AI
- Jonas Geiping 외. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach. arXiv:2502.05171 — 분야: cs.LG, cs.CL
- Kevin Lin 외. Sleep-time Compute: Beyond Inference Scaling at Test-time. arXiv:2504.13171 — 분야: cs.AI, cs.CL
- Yuxi Liu 외. The Serial Scaling Hypothesis. arXiv:2507.12549 — 분야: cs.LG, cs.CC, stat.ML
- Jingcheng Hu 외. PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning. arXiv:2601.05593 — 분야: cs.LG
[2026-05-27] 모델을 키우는 시대에서 하니스를 키우는 시대로 — 어제 그제의 두 글이 사실은 같은 분해의 사례였다
- 중심: Shangding Gu. From Model Scaling to System Scaling: Scaling the Harness in Agentic AI. arXiv:2605.26112 — 분야: cs.AI, cs.LG
- Romain Froger 외. ARE: Scaling Up Agent Environments and Evaluations. arXiv:2509.17158 — 분야: cs.AI, cs.CL
- Chingkwun Lam 외. Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework. arXiv:2603.11768 — 분야: cs.AI
- Aaditya Khanal 외. Beyond pass@1: A Reliability Science Framework for Long-Horizon LLM Agents. arXiv:2603.29231 — 분야: cs.AI
- Benjamin Rombaut. Inside the Scaffold: A Source-Code Taxonomy of Coding Agent Architectures. arXiv:2604.03515 — 분야: cs.SE, cs.AI, cs.ET
- Chenyu Zhou 외. Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering. arXiv:2604.08224 — 분야: cs.SE, cs.MA
- Jiahang Lin 외. Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses. arXiv:2604.25850 — 분야: cs.CL, cs.SE
- Hanxiang Chao 외. STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?. arXiv:2605.06527 — 분야: cs.CL
[2026-05-26] 확률과 결정론 사이의 이음새 — 어제의 로그가 정확히 어디서 갈라지는가
- 중심: Vasundra Srinivasan. A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents. arXiv:2605.20173 — 분야: cs.AI, cs.SE
- Matthew Thompson. The Dual-State Architecture for Reliable LLM Agents. arXiv:2512.20660 — 분야: cs.LG, cs.AI, cs.SE
- Raffi Khatchadourian. Replayable Financial Agents: A Determinism-Faithfulness Assurance Harness for Tool-Using LLM Agents. arXiv:2601.15322 — 분야: cs.AI, cs.CL
- Khush Patel 외. The Six Sigma Agent: Achieving Enterprise-Grade Reliability in LLM Systems Through Consensus-Driven Decomposed Execution. arXiv:2601.22290 — 분야: cs.AI
- Stephan Rabanser 외. Towards a Science of AI Agent Reliability. arXiv:2602.16666 — 분야: cs.AI, cs.CY, cs.LG
- Elzo Brito dos Santos Filho. ESAA: Event Sourcing for Autonomous Agents in LLM-Based Software Engineering. arXiv:2602.23193 — 분야: cs.AI
[2026-05-25] 로그가 곧 에이전트다 — 상태를 쌓지 말고 이벤트를 재투영하라
- 중심: Yohei Nakajima. The Log is the Agent: Event-Sourced Reactive Graphs for Auditable, Forkable Agentic Systems. arXiv:2605.21997 — 분야: cs.AI, cs.MA
- Erhu Feng 외. Get Experience from Practice: LLM Agents with Record & Replay. arXiv:2505.17716 — 분야: cs.LG, cs.MA
- Raffi Khatchadourian. Replayable Financial Agents: A Determinism-Faithfulness Assurance Harness for Tool-Using LLM Agents. arXiv:2601.15322 — 분야: cs.AI, cs.CL
- Elzo Brito dos Santos Filho. ESAA: Event Sourcing for Autonomous Agents in LLM-Based Software Engineering. arXiv:2602.23193 — 분야: cs.AI
- Yi Nian 외. Auditable Agents. arXiv:2604.05485 — 분야: cs.AI
- Josh Rosen, Seth Rosen. From Agent Loops to Deterministic Graphs: Execution Lineage for Reproducible AI-Native Work. arXiv:2605.06365 — 분야: cs.AI, cs.MA, cs.SE
[2026-05-24] SKILL.md는 수동 문서가 아니다 — 자연어만으로 레지스트리를 조작하는 의미적 공급망 공격
- 중심: Shoumik Saha 외. Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry. arXiv:2605.11418 — 분야: cs.AI, cs.CR
- Jonathan Sneh 외. ToolTweak: An Attack on Tool Selection in LLM-based Agents. arXiv:2510.02554 — 분야: cs.CR, cs.AI
- Yigitcan Kaya 외. When AI Meets the Web: Prompt Injection Risks in Third-Party AI Chatbot Plugins. arXiv:2511.05797 — 분야: cs.CR, cs.AI
- Narek Maloyan, Dmitry Namiot. Prompt Injection Attacks on Agentic Coding Assistants: A Systematic Analysis of Vulnerabilities in Skills, Tools, and Protocol Ecosystems. arXiv:2601.17548 — 분야: cs.CR
- Yi Liu 외. “Do Not Mention This to the User”: Detecting and Understanding Malicious Agent Skills in the Wild. arXiv:2602.06547 — 분야: cs.CR, cs.AI, cs.CL, cs.ET
- Zhiyuan Li 외. Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis. arXiv:2604.02837 — 분야: cs.CR, cs.AI
- Zenghao Duan 외. SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement. arXiv:2604.04989 — 분야: cs.CR
[2026-05-23] 측정을 측정하기 — 평가가 설계 과학이 되지 않으면 남는 것은 숫자뿐이다
- 중심: Keyang Xuan 외. Interactive Evaluation Requires a Design Science. arXiv:2605.17829 — 분야: cs.AI
- Kiana Jafari Meimandi 외. The Measurement Imbalance in Agentic AI Evaluation Undermines Industry Productivity Claims. arXiv:2506.02064 — 분야: cs.CY, cs.HC
- Victor Barres 외. $τ^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment. arXiv:2506.07982 — 분야: cs.AI, cs.CL
- Sri Vatsa Vuddanti, Satwik Kumar Chittiprolu. Recoverability Has a Law: The ERR Measure for Tool-Augmented Agents. arXiv:2601.22352 — 분야: cs.LG, cs.AI
- Yu Li 외. ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis. arXiv:2604.02022 — 분야: cs.AI
- Xiangyi Li 외. ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces. arXiv:2604.05172 — 분야: cs.AI
- Christopher Koch, Joshua Andreas Wellbrock. Beyond Task Success: An Evidence-Synthesis Framework for Evaluating, Governing, and Orchestrating Agentic AI. arXiv:2604.19818 — 분야: cs.SE, cs.HC, cs.MA
- Hao Wang 외. Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack. arXiv:2605.12673 — 분야: cs.AI, cs.CR
- Jiawei He 외. ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents. arXiv:2605.20251 — 분야: cs.SE, cs.AI
[2026-05-22] 기억이 가시권에 있어도 권위는 없다 — 암묵적 무효화와 쓰기측 판결
- 중심: Hanxiang Chao 외. STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?. arXiv:2605.06527 — 분야: cs.CL
- Yuanzhe Hu 외. Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions. arXiv:2507.05257 — 분야: cs.CL, cs.AI
- Xianda Zheng 외. Disentangling Reasoning Logic to Resolve Explicit Knowledge Conflicts. arXiv:2508.01273 — 분야: cs.AI
- Miao Su 외. Beyond Dialogue Time: Temporal Semantic Memory for Personalized LLM Agents. arXiv:2601.07468 — 분야: cs.AI
- Yiyang Feng 외. Tracking the Limits of Knowledge Propagation: How LLMs Fail at Multi-Step Reasoning with Conflicting Knowledge. arXiv:2601.15495 — 분야: cs.AI, cs.CL
- Xiaohui Zhang 외. ActMem: Bridging the Gap Between Memory Retrieval and Reasoning in LLM Agents. arXiv:2603.00026 — 분야: cs.CL, cs.AI, cs.IR
- Chingkwun Lam 외. Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework. arXiv:2603.11768 — 분야: cs.AI
- Ahmed Nusayer Ashik 외. When LLMs Lag Behind: Knowledge Conflicts from Evolving APIs in Code Generation. arXiv:2604.09515 — 분야: cs.SE
- Md Nayem Uddin 외. From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents. arXiv:2604.20006 — 분야: cs.CL
[2026-05-21] 유용한 기억이 망가질 때 — Consolidation 절차가 만드는 비단조적 붕괴
- 중심: Dylan Zhang 외. Useful Memories Become Faulty When Continuously Updated by LLMs. arXiv:2605.12978 — 분야: cs.AI
- Joon Sung Park 외. Generative Agents: Interactive Simulacra of Human Behavior. arXiv:2304.03442 — 분야: cs.HC, cs.AI, cs.LG
- Parth Sarthi 외. RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval. arXiv:2401.18059 — 분야: cs.CL, cs.LG
- Alex Laitenberger 외. Stronger Baselines for Retrieval-Augmented Generation with Long-Context Language Models. arXiv:2506.03989 — 분야: cs.CL
- Dongming Jiang 외. Anatomy of Agentic Memory: Taxonomy and Empirical Analysis of Evaluation and System Limitations. arXiv:2602.19320 — 분야: cs.CL, cs.AI
- Chingkwun Lam 외. Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework. arXiv:2603.11768 — 분야: cs.AI
- Jeonghye Kim 외. Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?. arXiv:2603.24472 — 분야: cs.CL, cs.LG
- Shu Wang 외. MemMachine: A Ground-Truth-Preserving Memory System for Personalized AI Agents. arXiv:2604.04853 — 분야: cs.AI
- Binyan Xu 외. Contextual Agentic Memory is a Memo, Not True Memory. arXiv:2604.27707 — 분야: cs.AI, cs.CL
- Hanxiang Chao 외. STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?. arXiv:2605.06527 — 분야: cs.CL
[2026-05-20] 상상 속에서 정책을 훈련한다는 것 — 마찰 우회의 두 번째 얼굴
- 중심: Nadav Timor 외. On Training in Imagination. arXiv:2605.06732 — 분야: cs.LG
- Leo Gao 외. Scaling Laws for Reward Model Overoptimization. arXiv:2210.10760 — 분야: cs.LG, stat.ML
- Emiliyan Gospodinov 외. Adaptive World Models: Learning Behaviors by Latent Imagination Under Non-Stationarity. arXiv:2411.01342 — 분야: cs.LG, cs.AI
- Jiawei Huang 외. Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage Perspective. arXiv:2502.19255 — 분야: cs.LG, cs.AI, stat.ML
- Mido Assran 외. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning. arXiv:2506.09985 — 분야: cs.AI, cs.CV, cs.LG, cs.RO
- Danijar Hafner 외. Training Agents Inside of Scalable World Models. arXiv:2509.24527 — 분야: cs.AI, cs.LG, cs.RO, stat.ML
- Zhennan Jiang 외. WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL. arXiv:2602.13977 — 분야: cs.RO, cs.AI
[2026-05-19] AI가 AI 연구자를 우회할 때 — 25명의 인터뷰가 드러낸 인식론적 분열
- 중심: Severin Field 외. AI Researchers’ Views on Automating AI R&D and Intelligence Explosions. arXiv:2603.03338 — 분야: cs.CY
- Severin Field. Why do Experts Disagree on Existential Risk and P(doom)? A Survey of AI Experts. arXiv:2502.14870 — 분야: cs.CY, cs.AI, cs.HC
- Joshua Clymer 외. Bare Minimum Mitigations for Autonomous AI Development. arXiv:2504.15416 — 분야: cs.CY
- Ning Li. The Ideation Bottleneck: Decomposing the Quality Gap Between AI-Generated and Human Economics Research. arXiv:2604.03338 — 분야: econ.GN, cs.AI, cs.CY
[2026-05-18] 스킬의 침식 — AI에 순응하는 인간이 잃는 것은 답이 아니라 오류와 씨름할 기회다
- 중심: Judy Hanwen Shen, Alex Tamkin. How AI Impacts Skill Formation. arXiv:2601.20245 — 분야: cs.CY, cs.AI, cs.HC
- Benjamin Lira 외. Coach not crutch: Evidence that AI can improve writing skill despite reducing effort. arXiv:2502.02880 — 분야: cs.HC
- Joel Becker 외. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv:2507.09089 — 분야: cs.AI, cs.HC, cs.SE
- Ali Aouad 외. Human-AI Productivity Paradoxes: Modeling the Interplay of Skill, Effort, and AI Assistance. arXiv:2605.11350 — 분야: cs.GT, cs.AI, econ.TH
[2026-05-17] 합의의 붕괴 — 다원성은 분포가 아니라 대화에서 살거나 죽는다
- 중심: Varad Vishwarupe 외. From Sycophantic Consensus to Pluralistic Repair: Why AI Alignment Must Surface Disagreement. arXiv:2605.14912 — 분야: cs.AI, cs.CY, cs.HC, cs.LG
- Mrinank Sharma 외. Towards Understanding Sycophancy in Language Models. arXiv:2310.13548 — 분야: cs.CL, cs.AI, cs.LG, stat.ML
- Taylor Sorensen 외. A Roadmap to Pluralistic Alignment. arXiv:2402.05070 — 분야: cs.AI, cs.CL, cs.IR
- Melody Y. Guan 외. Deliberative Alignment: Reasoning Enables Safer Language Models. arXiv:2412.16339 — 분야: cs.CL, cs.AI, cs.CY, cs.LG
- Jiseung Hong 외. Measuring Sycophancy of Language Models in Multi-turn Dialogues. arXiv:2505.23840 — 분야: cs.CL
- Huixin Zhong 외. Disentangling the Drivers of LLM Social Conformity: An Uncertainty-Moderated Dual-Process Mechanism. arXiv:2508.14918 — 분야: cs.CY, cs.AI
- Daniel Vennemeyer 외. Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs. arXiv:2509.21305 — 분야: cs.CL
- Liwei Jiang 외. Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond). arXiv:2510.22954 — 분야: cs.CL
- Itai Shapira 외. How RLHF Amplifies Sycophancy. arXiv:2602.01002 — 분야: cs.AI
- Kartik Chandra 외. Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians. arXiv:2602.19141 — 분야: cs.AI, cs.CY, cs.HC
[2026-05-16] 맥락 순응 — 검색이 틀렸을 때 RAG는 그것을 아는가
- 중심: Yihang Chen 외. Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict. arXiv:2605.14473 — 분야: cs.CL, cs.AI
- Shi-Qi Yan 외. RPO: Retrieval Preference Optimization for Robust Retrieval-Augmented Generation. arXiv:2501.13726 — 분야: cs.CL
- Chenyu Lin 외. Resisting Contextual Interference in RAG via Parametric-Knowledge Reinforcement. arXiv:2506.05154 — 분야: cs.CL, cs.AI, cs.IR
- Huixin Zhong 외. Disentangling the Drivers of LLM Social Conformity: An Uncertainty-Moderated Dual-Process Mechanism. arXiv:2508.14918 — 분야: cs.CY, cs.AI
- Yufeng Du 외. Context Length Alone Hurts LLM Performance Despite Perfect Retrieval. arXiv:2510.05381 — 분야: cs.CL, cs.AI
- Shuaizhi Cheng 외. The Override Gap: A Magnitude Account of Knowledge Conflict Failure in Hypernetwork-Based Instant LLM Adaptation. arXiv:2604.23750 — 분야: cs.LG, cs.AI
[2026-05-15] 방관자 효과 — 동료가 많아질수록 스스로 사고하기를 멈추는 LLM
- 중심: Dahlia Shehata, Ming Li. The Bystander Effect in Multi-Agent Reasoning: Quantifying Cognitive Loafing in Collaborative Interactions. arXiv:2605.10698 — 분야: cs.MA, cs.AI
- Lin Shi 외. Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge. arXiv:2406.07791 — 분야: cs.CL, cs.AI
- Wenzhe Li 외. Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?. arXiv:2502.00674 — 분야: cs.CL, cs.LG
- Yuxuan Li 외. Systematic Failures in Collective Reasoning under Distributed Information in Multi-Agent LLMs. arXiv:2505.11556 — 분야: cs.CL, cs.AI, cs.MA
- Keyu Wang 외. When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models. arXiv:2508.02087 — 분야: cs.CL
- Zhiwei Zhang 외. Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation. arXiv:2511.02303 — 분야: cs.AI, cs.CL
[2026-05-14] 메모리 저주 — 더 많이 기억할수록 덜 협동하는 LLM
- 중심: Jiayuan Liu 외. The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents. arXiv:2605.08060 — 분야: cs.CL, cs.AI, cs.GT, cs.MA
- Jingru Jia 외. LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory. arXiv:2502.20432 — 분야: cs.AI, cs.CY, cs.GT, cs.LG
- Taisei Hishiki 외. How memory can affect collective and cooperative behaviors in an LLM-Based Social Particle Swarm. arXiv:2604.12250 — 분야: cs.AI, cs.CL, cs.GT, cs.MA
- Emanuel Tewolde 외. CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas. arXiv:2604.15267 — 분야: cs.GT, cs.AI, cs.CL, cs.CY, cs.MA
[2026-05-13] 토큰이 자신을 잊지 않으려면 — TIDE와 레이어마다 되새기는 정체성
- 중심: Ajay Jaiswal 외. TIDE: Every Layer Knows the Token Beneath the Context. arXiv:2605.06216 — 분야: cs.CL, cs.AI, cs.LG
- Da Yu 외. Scaling Embedding Layers in Language Models. arXiv:2502.01637 — 분야: cs.CL, cs.LG
- Jing Liu 외. Distributed Specialization: Rare-Token Neurons in Large Language Models. arXiv:2509.21163 — 분야: cs.AI
- Hong Liu 외. Scaling Embeddings Outperforms Scaling Experts in Language Models. arXiv:2601.21204 — 분야: cs.CL, cs.AI, cs.LG
[2026-05-10] RL이 가르칠 수 있는 것의 모양 — 표현성이 멱법칙을 어떻게 휘게 하는가
- 중심: Tianle Wang 외. Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key. arXiv:2605.06638 — 분야: cs.AI, cs.CL
- Yang Yue 외. Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?. arXiv:2504.13837 — 분야: cs.AI, cs.CL, cs.CV
- Zelin Tan 외. Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning. arXiv:2509.25300 — 분야: cs.LG, cs.AI
- Sunghwan Kim 외. On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length. arXiv:2605.02572 — 분야: cs.AI, cs.LG
- Ömer Faruk Akgül 외. Rethinking RL for LLM Reasoning: It’s Sparse Policy Selection, Not Capability Learning. arXiv:2605.06241 — 분야: cs.CL
[2026-05-05] 단어 없이 생각하기 — 64개 추상 토큰이 만드는 이산 잠재 추론
- 중심: Keshav Ramji 외. Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought. arXiv:2604.22709 — 분야: cs.CL
- Aaron van den Oord 외. Neural Discrete Representation Learning. arXiv:1711.00937 — 분야: cs.LG
- Shibo Hao 외. Training Large Language Models to Reason in a Continuous Latent Space. arXiv:2412.06769 — 분야: cs.CL
- DiJia Su 외. Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning. arXiv:2502.03275 — 분야: cs.CL, cs.AI, cs.LG, cs.LO
- Zhenyi Shen 외. CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation. arXiv:2502.21074 — 분야: cs.CL
- Jingxian Xu 외. TwT: Thinking without Tokens by Habitual Reasoning Distillation with Multi-Teachers’ Guidance. arXiv:2503.24198 — 분야: cs.CL
- Zehong Wang 외. Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents. arXiv:2601.22311 — 분야: cs.AI, cs.CL, cs.LG
- Jiaxuan Zou 외. Capabilities and Fundamental Limits of Latent Chain-of-Thought. arXiv:2602.01148 — 분야: cs.AI, cs.IT, cs.LG, math.OC
- Wenshuo Wang. LLM Reasoning Is Latent, Not the Chain of Thought. arXiv:2604.15726 — 분야: cs.AI
- Yuyan Zhou 외. LEPO: Latent Reasoning Policy Optimization for Large Language Models. arXiv:2604.17892 — 분야: cs.LG, cs.AI
[2026-05-03] 재귀로 묶인 다중 에이전트 — 잠재공간이 텍스트 병목을 우회할 때
- 중심: Xiyuan Yang 외. Recursive Multi-Agent Systems. arXiv:2604.25917 — 분야: cs.AI, cs.CL, cs.LG
- Jakob N. Foerster 외. Learning to Communicate with Deep Multi-Agent Reinforcement Learning. arXiv:1605.06676 — 분야: cs.AI, cs.LG, cs.MA
- Sainbayar Sukhbaatar 외. Learning Multiagent Communication with Backpropagation. arXiv:1605.07736 — 분야: cs.LG, cs.AI
- Mostafa Dehghani 외. Universal Transformers. arXiv:1807.03819 — 분야: cs.CL, cs.LG, stat.ML
- Zhenzhong Lan 외. ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. arXiv:1909.11942 — 분야: cs.CL, cs.AI
- Evan Hubinger 외. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training. arXiv:2401.05566 — 분야: cs.CR, cs.AI, cs.CL, cs.LG, cs.SE
- Minyoung Huh 외. The Platonic Representation Hypothesis. arXiv:2405.07987 — 분야: cs.LG, cs.AI, cs.CV, cs.NE
- Shibo Hao 외. Training Large Language Models to Reason in a Continuous Latent Space. arXiv:2412.06769 — 분야: cs.CL
- Yuichi Inoue 외. Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree Search. arXiv:2503.04412 — 분야: cs.AI
- Xin Wei Chia 외. Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States. arXiv:2503.09066 — 분야: cs.LG, cs.AI, cs.CR
- Zhexuan Wang 외. AgentDropout: Dynamic Agent Elimination for Token-Efficient and High-Performance LLM-Based Multi-Agent Collaboration. arXiv:2503.18891 — 분야: cs.CL, cs.AI
- Rui-Jie Zhu 외. Scaling Latent Reasoning via Looped Language Models. arXiv:2510.25741 — 분야: cs.CL
- Fu-Chun Yang, Jason Eshraghian. Direct Semantic Communication Between Large Language Models via Vector Translation. arXiv:2511.03945 — 분야: cs.CL, cs.AI
- Zhuoyun Du 외. Enabling Agents to Communicate Entirely in Latent Space. arXiv:2511.09149 — 분야: cs.LG, cs.AI, cs.MA
- Jiaru Zou 외. Latent Collaboration in Multi-Agent Systems. arXiv:2511.20639 — 분야: cs.CL, cs.AI, cs.LG
- Hayden Prairie 외. Parcae: Scaling Laws For Stable Looped Language Models. arXiv:2604.12946 — 분야: cs.LG
[2026-05-02] 표면 아래의 LLM — 문해는 늘었지만 함의는 못 짓는다
- 중심: Kabir Ahuja 외. Beneath the Surface: Investigating LLMs’ Capabilities for Communicating with Subtext. arXiv:2604.05273 — 분야: cs.CL
- Omar Shaikh 외. Grounding Gaps in Language Model Generations. arXiv:2311.09144 — 분야: cs.CL, cs.HC
- Joshua Tint 외. ExpressivityBench: Can LLMs Communicate Implicitly?. arXiv:2411.08010 — 분야: cs.CL, cs.AI
- Joshua Lee 외. Pragmatic Metacognitive Prompting Improves LLM Performance on Sarcasm Detection. arXiv:2412.04509 — 분야: cs.CL
- Kefan Yu 외. The Pragmatic Mind of Machines: Tracing the Emergence of Pragmatic Competence in Large Language Models. arXiv:2505.18497 — 분야: cs.CL
- Saki Imai 외. Measuring How (Not Just Whether) VLMs Build Common Ground. arXiv:2509.03805 — 분야: cs.CL, cs.AI
- Takuma Sato 외. Pragmatic Theories Enhance Understanding of Implied Meanings in LLMs. arXiv:2510.26253 — 분야: cs.CL
- Christian Nickel 외. Understanding Artificial Theory of Mind: Perturbed Tasks and Reasoning in Large Language Models. arXiv:2602.22072 — 분야: cs.CL, cs.AI
- Ruirui Chen 외. CoMMET: To What Extent Can LLMs Perform Theory of Mind Tasks?. arXiv:2603.11915 — 분야: cs.CL
- Guangsheng Yu, Xu Wang. Knows: Agent-Native Structured Research Representations. arXiv:2604.17309 — 분야: cs.AI
- Xiyuan Yang 외. Recursive Multi-Agent Systems. arXiv:2604.25917 — 분야: cs.AI, cs.CL, cs.LG
[2026-05-01] 마지막 사람-쓴 논문 — 두 가지 세금과 ARA의 약속, 그리고 족쇄
- 중심: Jiachen Liu 외. The Last Human-Written Paper: Agent-Native Research Artifacts. arXiv:2604.24658 — 분야: cs.LG
- Yufeng Du 외. Context Length Alone Hurts LLM Performance Despite Perfect Retrieval. arXiv:2510.05381 — 분야: cs.CL, cs.AI
- Kabir Ahuja 외. Beneath the Surface: Investigating LLMs’ Capabilities for Communicating with Subtext. arXiv:2604.05273 — 분야: cs.CL
- Guangsheng Yu, Xu Wang. Knows: Agent-Native Structured Research Representations. arXiv:2604.17309 — 분야: cs.AI
- Xiyuan Yang 외. Recursive Multi-Agent Systems. arXiv:2604.25917 — 분야: cs.AI, cs.CL, cs.LG
[2026-04-30] MCP의 도구세 — Tool Attention이 제안한 해법과 그 한계
- 중심: Anuj Sadani, Deepak Kumar. Tool Attention Is All You Need: Dynamic Tool Gating and Lazy Schema Loading for Eliminating the MCP/Tools Tax in Scalable Agentic Workflows. arXiv:2604.21816 — 분야: cs.AI
- Ahilan Ayyachamy Nadar Ponnusamy 외. Context Discipline and Performance Correlation: Analyzing LLM Performance and Quality Degradation Under Varying Context Lengths. arXiv:2601.11564 — 분야: cs.CL, cs.AI
- Mohammed Mehedi Hasan 외. Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions. arXiv:2602.14878 — 분야: cs.SE, cs.ET
- Uria Franko. Dynamic System Instructions and Tool Exposure for Efficient Agentic LLMs. arXiv:2602.17046 — 분야: cs.AI
- Charoes Huang 외. Model Context Protocol Threat Modeling and Analyzing Vulnerabilities to Prompt Injection with Tool Poisoning. arXiv:2603.22489 — 분야: cs.CR, cs.SE
[2026-04-29] 웹 에이전트의 계획 — 탐색 알고리즘으로 다시 본 LLM 행위자
- 중심: Orit Shahnovsky, Rotem Dror. AI Planning Framework for LLM-Based Web Agents. arXiv:2603.12710 — 분야: cs.AI, cs.CL
- Xing Han Lù 외. AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories. arXiv:2504.08942 — 분야: cs.LG, cs.AI, cs.CL
- Davide Paglieri 외. Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents. arXiv:2509.03581 — 분야: cs.AI
- Yanyu Chen 외. TRACE: Trajectory-Aware Comprehensive Evaluation for Deep Research Agents. arXiv:2602.21230 — 분야: cs.CL
- Mohamed Aghzal 외. Why Do LLM-based Web Agents Fail? A Hierarchical Planning Perspective. arXiv:2603.14248 — 분야: cs.AI, cs.CL
[2026-04-28] 자기 자신을 편집하는 모델 — MEMENTO가 보여준 것과 포기한 것
- 중심: Vasilis Kontonis 외. MEMENTO: Teaching LLMs to Manage Their Own Context. arXiv:2604.09852 — 분야: cs.AI, cs.LG
[2026-04-27] 메모리를 비우니 감사 가능성이 보였다 — DPM이 RAG의 진짜 이유를 짚다
- 중심: Vasundra Srinivasan. Stateless Decision Memory for Enterprise AI Agents. arXiv:2604.20158 — 분야: cs.AI
[2026-04-26] 플랫 메모리의 맹점 — StructMem이 짚어낸 것
- 중심: Buqiang Xu 외. StructMem: Structured Memory for Long-Horizon Behavior in LLMs. arXiv:2604.21748 — 분야: cs.CL, cs.AI, cs.IR, cs.LG, cs.MA
[2026-04-25] 모델 안의 사회 — RL이 스스로 발견한 다관점 대화
- Yingxuan Yang 외. Understanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity. arXiv:2602.03794 — 분야: cs.AI, cs.LG
- James Evans 외. Agentic AI and the next intelligence explosion. arXiv:2603.20639 — 분야: cs.AI
[2026-04-25] 재귀의 안쪽 — 우리 작업 자체가 multi-agent system인 이유
- James Evans 외. Agentic AI and the next intelligence explosion. arXiv:2603.20639 — 분야: cs.AI
[2026-04-25] 고무 도장 심판, 숨겨진 프로파일 — 거버넌스 실패가 공학 실험에 나타나는 방식
- Yubin Kim 외. Towards a Science of Scaling Agent Systems. arXiv:2512.08296 — 분야: cs.AI
[2026-04-23] Aggregator, Planner, Manager — 다른 이름, 같은 자리
- Junlin Wang 외. Mixture-of-Agents Enhances Large Language Model Capabilities. arXiv:2406.04692 — 분야: cs.CL
- Yubin Kim 외. Towards a Science of Scaling Agent Systems. arXiv:2512.08296 — 분야: cs.AI
[2026-04-21] 에이전트를 더 넣으면 왜 나아지지 않는가 — 상한과 하한의 공존
- Yubin Kim 외. Towards a Science of Scaling Agent Systems. arXiv:2512.08296 — 분야: cs.AI
- Yingxuan Yang 외. Understanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity. arXiv:2602.03794 — 분야: cs.AI, cs.LG