Monthly ranking

2026-06

Top 10 papers, rising papers, dominant themes, and the full monthly index.

Top 10 Papers

  1. Agentic Abstention: Do Agents Know When to Stop Instead of Act?100
  2. Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent100
  3. The Verification Horizon: No Silver Bullet for Coding Agent Rewards100
  4. TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents99
  5. DanceOPD: On-Policy Generative Field Distillation98
  6. OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning98
  7. Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction97
  8. JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting97
  9. ReFreeKV: Towards Threshold-Free KV Cache Compression96
  10. Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents95

Themes

  • reinforcement learning4
  • diffusion models3
  • foundation models3
  • large language models3
  • GRPO2
  • GUI agents2

Complete Monthly Index

  1. Agentic Abstention: Do Agents Know When to Stop Instead of Act?CONVOLVE, LLM-as-agent systems, agentic abstention, context engineering, question answering, sequential decision problem
  2. Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B AgentMixture-of-Experts, agent horizon, agentic model, agentic trajectories, domain-level teacher models, domain-routed training
  3. The Verification Horizon: No Silver Bullet for Coding Agent Rewardsgenerative capabilities, human intent, policy capability, proxy signals, reward design, reward hacking
  4. TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agentsbenchmark evaluation, computer-use tasks, digital activities, execution-based scoring protocol, general-purpose agents, graphical user interfaces
  5. DanceOPD: On-Policy Generative Field Distillationclassifier-free guidance, expert capabilities, flow-matching models, generative field distillation, global editing, local editing
  6. OPID: On-Policy Skill Distillation for Agentic Reinforcement Learningcritical-first routing, hierarchical skills, on-policy trajectories, outcome-based reinforcement learning, policy optimization, reinforcement learning
  7. Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe ExtractionGUI agents, Multimodal Large Language Models, Video Question Answering, keyframe extraction, scene dynamics, task relevance
  8. JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree DraftingMoE Qwen3, acceptance rate, autoregressive Large Language Models, autoregressive factorization, bidirectional block-diffusion, branch-agnostic marginals
  9. ReFreeKV: Towards Threshold-Free KV Cache CompressionKV cache compression, KV cache pruning, LLM inference, adaptive budget allocation, threshold-free methods
  10. Neglected Free Lunch from Post-training: Progress Advantage for LLM AgentsMarkov decision process, advantage function, agentic settings, failure attribution, log-probability ratio, progress advantage
  11. OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasksagent-pattern challenges, binary-completion metric, computer-use workflows, cross-source reasoning, implicit-state inference, long-horizon tasks
  12. SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoningcross-modal joint-risk, dynamic-rule evaluation, fast--slow decoupled reinforcement learning, multimodal conversations, multimodal guardrail benchmark, multimodal guardrail model
  13. TACO: Tool-Augmented Credit Optimization for Agentic Tool UseDifferential Answer-Probe Reward, GRPO, Outcome-Gated Advantage Routing, SFT+RL pipeline, agentic multimodal models, answer checker
  14. Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMsLLMs, axiomatic evaluation framework, causality, downstream benchmark scores, factual QA, functional axioms
  15. Monte Carlo Energy Aggregation for Mobile 3D Gaussian SplattingAttribute-Conditioned SH Enhancement, Gaussian Splatting, Monte Carlo Specular Energy Aggregator, Spherical Harmonics, latent space, mobile platforms
  16. Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generationagentic framework, context gap, context grounding, context-aware planning, image agent bench, image agent capabilities
  17. ViQ: Text-Aligned Visual Quantized Representations at Any Resolutiondiscrete representations, feature discretization, low-level reconstruction, multimodal modeling, position-aware head-wise quantization, proximal representation learning
  18. Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) BenchmarkNanotechnology Molecular Optimization, domain-agnostic pretraining, fitness landscapes, generative models, molecular optimization, pharmaceutical dataset bias
  19. GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated ScreenshotsGUI agents, GUI interaction, curriculum learning, reinforcement learning, screen shots, visual grounding
  20. ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question AnsweringKnowledge-based Visual Question Answering, TN-GSPO, deduplication, generation length, multimodal search agent, rejection-sampling SFT
  21. AsyncOPD: How Stale Can On-Policy Distillation Be?KL divergence, Monte Carlo estimation, asynchronous training, forward KL, on-policy distillation, policy gradient
  22. Boundary-Aware Context Grounding for A Low-Channel EEG Agentboundary awareness, context pack, electroencephalography, hardware-aware grounding, large language models, machine-readable artifacts
  23. GBC: Gradient-Based Connections for Optimizing Multi-Agent Systemsagent coordination, attribution graph, computational graph, credit assignment, gradient-based connection weights, large language models
  24. PhysisForcing: Physics Reinforced World Simulator for Robotic ManipulationDiT features, EZS-Bench, PAI-Bench, R-Bench, WorldArena, action-planner protocol
  25. ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and GenerationGRPO, boundary-aware count policy, count-faithful image generation, crowd counting, cycle-consistent learning, density-aware adaptive zooming
  26. COrigami: An AI Pipeline for Co-Designing Flat-Foldable Visually Recognisable Origamiaesthetic evaluation, base packing, co-creativity, computational origami, crease patterns, flat foldability
  27. DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Modelautoregressive video stack, consumer-GPU runtime, dual-view operation, interactive rollouts, mid-stream reprompting, multimodal initialization
  28. Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)AWR, DAgger-like HIL, HuggingFace Hub, RECAP, Thompson sampling, advantage estimation
  29. TheoremGraph: Bridging Formal and Informal MathematicsLLM judge, LeanGraph, LeanSearch, TheoremGraph, concept retrieval, formal libraries
  30. Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical ReasoningVideo-MME-Logical, dynamic spatiality, multimodal large language models, sequential counting, state tracking, structural composition
  31. GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use AgentsN/A
  32. Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image GenerationILScore, free-length multimodal token sequences, interleaved text-image sequences, multimodal data efficiency, multimodal intelligence, multimodal training process
  33. Large-Scale Tunnel Air-Ground Collaboration With FLISP: Fast LiDAR-IMU Synchronized Path PlannerN/A
  34. PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM AgentsLLM agents, argument-level guards, conversation context, dialogue reasoning, policy adherence, policy violation recall
  35. SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailingguardrails, in-context policy guardrailing, multi-turn conversations, natural-language rules, policy frameworks, policy specifications
  36. SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluationaffordance-preserving variations, digital cousins, digital twins, policy training, real-to-sim scene construction, robotic manipulation
  37. Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice AgentsPareto frontier, conversational infill, foundation models, latency, millisecond-level time-to-first-response, real-time models
  38. ZooClaw-FashionSigLIP2: Distilled Fine-tuning for Robust Fashion RetrievalLoRA, SigLIP2-base, benchmark evaluation, fashion retrieval, full fine-tuning, ground truth
  39. Cluster, Route, Escalate: Cascaded Framework for Cost-Aware LLM ServingTime Per Output Token, cascaded solution, cost-effective model, large language models, model routing, quality estimation
  40. How Good Can Linear Models Be for Time-Series Forecasting?Ridge regression, augmentation, context length, cross-series sharing, forecast horizon, foundation models
  41. LiveEdit: Towards Real-Time Diffusion-Based Streaming Video EditingAR-oriented mask cache, augmented reality, bidirectional foundation model, causal editing, content preservation, frame-by-frame editing
  42. One Forward Beats Two: InnerZoom for Accurate and Efficient GUI GroundingMLLM-based GUI grounding, SFT+RL, TFLOPs, ZoomIn-style methods, autoregressive coordinate generation, cross-layer evidence bridging
  43. Trimming the Long-Tail of Visual World Modeling Evaluationaffordance generalization, constraint awareness, descriptive generation, image generation, impossible scenarios, long-tailed distribution
  44. Information-Aware KV Cache Compression for Long ReasoningForward Influence, KV cache, KV cache compression, LLMs, attention weights, entropy-aware
  45. Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image SynthesisDPG, GenEval, Grouped Cross-Entropy, HPSv3, VRAM usage, discrete tokens
  46. Towards Automating Scientific Review with Google's Paper Assistant ToolAI-assisted scientific discovery, AI-human collaboration, SPOT benchmark, agentic AI framework, inference scaling, mathematical errors
  47. When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier ModelsClopper-Pearson bound, GPQA-Diamond, Gaussian copula, Self-MoA, accuracy, beta
  48. Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix Itagentic reinforcement learning, catastrophic collapse, control tokens, erroneous example supervision, exploratory learning, hint-based guidance
  49. Walking in the Implicit: Interactive World Exploration via Neural Scene RepresentationNeural Implicit Scene, VAE encoder, camera trajectories, diffusion transformer, geometry-aware retrieval, implicit state
  50. How Post-Training Shapes Biological Reasoning Modelscontinued pre-training, foundation models, generalization, in-domain performance, language models, multimodal biological data
  51. LISA: Likelihood Score Alignment for Visual-condition Controllable GenerationLISA, conditional control, decoder, diffusion models, disentangled features, feature projection
  52. MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image GenerationFID, Masked Image Modeling, Normalizing Flows, VAE encoder, generative flow, high-frequency synthesis
  53. Discretizing Reward ModelsMonte Carlo dropout, discretization, discriminative ability, oversensitivity, policy learning, reinforcement learning
  54. EO-WM: A Physically Informed World Model for Probabilistic Earth Observation ForecastingNDVI, Normalized Difference Vegetation Index, climatological baseline, cumulative physical stress signals, diffusion models, meteorological forcing
  55. Hallucination in World Models is Predictable and Preventablecoverage-aware sampling, curiosity rewards, data-centric signals, data-efficient fine-tuning, ground-truth actions, hallucination
  56. Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA EnhancementVision-Language-Action models, domain gap, imitation learning, object-centric representation, pose estimation, reinforcement learning
  57. The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of TatarTatar language, comparative experiments, cross lingual transfer, evaluation, fine tuning, low resource languages
  58. RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation6D Plucker ray coordinates, Vision Foundation Models, any-resolution cross-attention, dense prediction tasks, feature upsampling, geometry-aware neighborhood attention
  59. CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent EconomiesLLM agents, agent behavior, autonomous agents, communication, cumulative net income, economic systems
  60. Ko-WideSearch: A Korean Breadth-Search Benchmark for Exhaustive Set Enumeration by Web AgentsColumn-F1, Item-F1, Row-F1, automated synthesize-and-verify pipeline, breadth-search, composite key
  61. OpenBioRQ: Unsolved Biomedical Research Questions for Agentsagentic collapse, agentic models, answer key, biomedical research questions, citation verification, frontier agents
  62. Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoEMixture-of-Experts, compute allocation, denoising process, diffusion models, latent features, post-training framework
  63. Interleaved Speech Language Models Latently Work In Textintermediate layers, logit lens, speech language models, speech recognition, speech-text interleaving, spoken knowledge abilities
  64. NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement LearningMLLM-judged image quality, adjoint sensitivity analysis, classifier-free guidance, flow-based generators, forensic realism, hinge penalty
  65. The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction4D hand motion reconstruction, egocentric video, full frames, hand-overlay rendering, hand-pose annotations, metric-scale pose
  66. Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robotsattention masking, bi-manual manipulation, embodiment differences, head-camera frame, interleaved action tokens, parallel grippers
  67. Learning Transferable Dynamics Priors from Action to World Modelingaction-conditioned, diffusion world model, dynamics priors, multi-view interactive, policy-centric learning, pretraining
  68. PoseShield: Neural Collision Fields for Human Self-Collision ResolutionEikonal equation, SMPL, constrained optimization, motion generation, neural collision constraint, pose estimation
  69. Fast LeWorldModelJoint-Embedding Predictive Architectures, LeWorldModel, action-prefix prediction, autoregressive rollout, latent transition model, latent world model
  70. Confidence-Aware Tool Orchestration for Robust Video UnderstandingBlind Trust Problem, agentic video understanding, calibrated reliability score, confidence-cost GRPO reward, evidence interface, reliability-relevance score
  71. Geometric Stability of Neural Population Codes: Regional Variation, Behavioral Relevance, and Circuit DependenceSpearman rank correlation, attractor network model, feedforward input, geometric stability, hippocampal circuits, pairwise distance structure
  72. Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments3D reasoning, agent generalization, agentic systems, automated evaluation engine, benchmark, graphical understanding
  73. Beyond IID: How General Are Tabular Foundation Models, Really?Data Foundry, IID data, benchmarking, deep learning models, high-dimensional datasets, non-IID data
  74. Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty PredictionLarge Reasoning Models, cognitive episodes, difficulty prediction, effort allocation, episode sequences, episode-dynamic features
  75. Qwen-Image-2.0-RL Technical ReportGRPO-based RL training framework, chain-of-thought reasoning, hybrid classifier-free guidance, image editing, intra-group reward range filtering, on-policy distillation
  76. MultiHashFormer: Hash-based Generative Language ModelsHash Decoder, Hash Encoder, MultiHashFormer, Transformer decoder, causal LMs, discrete hash IDs
  77. PhysiFormer: Learning to Simulate Mechanics in World Space3D meshes, attention factorised, autoregressive baselines, denoising diffusion process, diffusion transformer, permutation-invariant
  78. In-Context World Modeling for Robotic ControlVision-Language-Action models, in-context adaptation, novel configurations, parameter updates, real-world robot platforms, robot policies
  79. ReasoningLens: Hierarchical Visualization and Diagnostic Auditing for Large Reasoning ModelsChain-of-Thought traces, agentic auditor, diagnostic auditing, hierarchical visualization, information necropsy, systemic reasoning profiles