Monthly ranking
2026-06
Top 10 papers, rising papers, dominant themes, and the full monthly index.
Top 10 Papers
- Agentic Abstention: Do Agents Know When to Stop Instead of Act?100
- Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent100
- The Verification Horizon: No Silver Bullet for Coding Agent Rewards100
- TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents99
- DanceOPD: On-Policy Generative Field Distillation98
- OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning98
- Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction97
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting97
- ReFreeKV: Towards Threshold-Free KV Cache Compression96
- Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents95
Rising Papers
- Agentic Abstention: Do Agents Know When to Stop Instead of Act?100/100
- Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent100/100
- The Verification Horizon: No Silver Bullet for Coding Agent Rewards100/100
- TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents99/100
- DanceOPD: On-Policy Generative Field Distillation98/100
Themes
- reinforcement learning4
- diffusion models3
- foundation models3
- large language models3
- GRPO2
- GUI agents2
Complete Monthly Index
- Agentic Abstention: Do Agents Know When to Stop Instead of Act?CONVOLVE, LLM-as-agent systems, agentic abstention, context engineering, question answering, sequential decision problem
- Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B AgentMixture-of-Experts, agent horizon, agentic model, agentic trajectories, domain-level teacher models, domain-routed training
- The Verification Horizon: No Silver Bullet for Coding Agent Rewardsgenerative capabilities, human intent, policy capability, proxy signals, reward design, reward hacking
- TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agentsbenchmark evaluation, computer-use tasks, digital activities, execution-based scoring protocol, general-purpose agents, graphical user interfaces
- DanceOPD: On-Policy Generative Field Distillationclassifier-free guidance, expert capabilities, flow-matching models, generative field distillation, global editing, local editing
- OPID: On-Policy Skill Distillation for Agentic Reinforcement Learningcritical-first routing, hierarchical skills, on-policy trajectories, outcome-based reinforcement learning, policy optimization, reinforcement learning
- Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe ExtractionGUI agents, Multimodal Large Language Models, Video Question Answering, keyframe extraction, scene dynamics, task relevance
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree DraftingMoE Qwen3, acceptance rate, autoregressive Large Language Models, autoregressive factorization, bidirectional block-diffusion, branch-agnostic marginals
- ReFreeKV: Towards Threshold-Free KV Cache CompressionKV cache compression, KV cache pruning, LLM inference, adaptive budget allocation, threshold-free methods
- Neglected Free Lunch from Post-training: Progress Advantage for LLM AgentsMarkov decision process, advantage function, agentic settings, failure attribution, log-probability ratio, progress advantage
- OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasksagent-pattern challenges, binary-completion metric, computer-use workflows, cross-source reasoning, implicit-state inference, long-horizon tasks
- SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoningcross-modal joint-risk, dynamic-rule evaluation, fast--slow decoupled reinforcement learning, multimodal conversations, multimodal guardrail benchmark, multimodal guardrail model
- TACO: Tool-Augmented Credit Optimization for Agentic Tool UseDifferential Answer-Probe Reward, GRPO, Outcome-Gated Advantage Routing, SFT+RL pipeline, agentic multimodal models, answer checker
- Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMsLLMs, axiomatic evaluation framework, causality, downstream benchmark scores, factual QA, functional axioms
- Monte Carlo Energy Aggregation for Mobile 3D Gaussian SplattingAttribute-Conditioned SH Enhancement, Gaussian Splatting, Monte Carlo Specular Energy Aggregator, Spherical Harmonics, latent space, mobile platforms
- Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generationagentic framework, context gap, context grounding, context-aware planning, image agent bench, image agent capabilities
- ViQ: Text-Aligned Visual Quantized Representations at Any Resolutiondiscrete representations, feature discretization, low-level reconstruction, multimodal modeling, position-aware head-wise quantization, proximal representation learning
- Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) BenchmarkNanotechnology Molecular Optimization, domain-agnostic pretraining, fitness landscapes, generative models, molecular optimization, pharmaceutical dataset bias
- GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated ScreenshotsGUI agents, GUI interaction, curriculum learning, reinforcement learning, screen shots, visual grounding
- ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question AnsweringKnowledge-based Visual Question Answering, TN-GSPO, deduplication, generation length, multimodal search agent, rejection-sampling SFT
- AsyncOPD: How Stale Can On-Policy Distillation Be?KL divergence, Monte Carlo estimation, asynchronous training, forward KL, on-policy distillation, policy gradient
- Boundary-Aware Context Grounding for A Low-Channel EEG Agentboundary awareness, context pack, electroencephalography, hardware-aware grounding, large language models, machine-readable artifacts
- GBC: Gradient-Based Connections for Optimizing Multi-Agent Systemsagent coordination, attribution graph, computational graph, credit assignment, gradient-based connection weights, large language models
- PhysisForcing: Physics Reinforced World Simulator for Robotic ManipulationDiT features, EZS-Bench, PAI-Bench, R-Bench, WorldArena, action-planner protocol
- ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and GenerationGRPO, boundary-aware count policy, count-faithful image generation, crowd counting, cycle-consistent learning, density-aware adaptive zooming
- COrigami: An AI Pipeline for Co-Designing Flat-Foldable Visually Recognisable Origamiaesthetic evaluation, base packing, co-creativity, computational origami, crease patterns, flat foldability
- DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Modelautoregressive video stack, consumer-GPU runtime, dual-view operation, interactive rollouts, mid-stream reprompting, multimodal initialization
- Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)AWR, DAgger-like HIL, HuggingFace Hub, RECAP, Thompson sampling, advantage estimation
- TheoremGraph: Bridging Formal and Informal MathematicsLLM judge, LeanGraph, LeanSearch, TheoremGraph, concept retrieval, formal libraries
- Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical ReasoningVideo-MME-Logical, dynamic spatiality, multimodal large language models, sequential counting, state tracking, structural composition
- GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use AgentsN/A
- Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image GenerationILScore, free-length multimodal token sequences, interleaved text-image sequences, multimodal data efficiency, multimodal intelligence, multimodal training process
- Large-Scale Tunnel Air-Ground Collaboration With FLISP: Fast LiDAR-IMU Synchronized Path PlannerN/A
- PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM AgentsLLM agents, argument-level guards, conversation context, dialogue reasoning, policy adherence, policy violation recall
- SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailingguardrails, in-context policy guardrailing, multi-turn conversations, natural-language rules, policy frameworks, policy specifications
- SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluationaffordance-preserving variations, digital cousins, digital twins, policy training, real-to-sim scene construction, robotic manipulation
- Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice AgentsPareto frontier, conversational infill, foundation models, latency, millisecond-level time-to-first-response, real-time models
- ZooClaw-FashionSigLIP2: Distilled Fine-tuning for Robust Fashion RetrievalLoRA, SigLIP2-base, benchmark evaluation, fashion retrieval, full fine-tuning, ground truth
- Cluster, Route, Escalate: Cascaded Framework for Cost-Aware LLM ServingTime Per Output Token, cascaded solution, cost-effective model, large language models, model routing, quality estimation
- How Good Can Linear Models Be for Time-Series Forecasting?Ridge regression, augmentation, context length, cross-series sharing, forecast horizon, foundation models
- LiveEdit: Towards Real-Time Diffusion-Based Streaming Video EditingAR-oriented mask cache, augmented reality, bidirectional foundation model, causal editing, content preservation, frame-by-frame editing
- One Forward Beats Two: InnerZoom for Accurate and Efficient GUI GroundingMLLM-based GUI grounding, SFT+RL, TFLOPs, ZoomIn-style methods, autoregressive coordinate generation, cross-layer evidence bridging
- Trimming the Long-Tail of Visual World Modeling Evaluationaffordance generalization, constraint awareness, descriptive generation, image generation, impossible scenarios, long-tailed distribution
- Information-Aware KV Cache Compression for Long ReasoningForward Influence, KV cache, KV cache compression, LLMs, attention weights, entropy-aware
- Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image SynthesisDPG, GenEval, Grouped Cross-Entropy, HPSv3, VRAM usage, discrete tokens
- Towards Automating Scientific Review with Google's Paper Assistant ToolAI-assisted scientific discovery, AI-human collaboration, SPOT benchmark, agentic AI framework, inference scaling, mathematical errors
- When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier ModelsClopper-Pearson bound, GPQA-Diamond, Gaussian copula, Self-MoA, accuracy, beta
- Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix Itagentic reinforcement learning, catastrophic collapse, control tokens, erroneous example supervision, exploratory learning, hint-based guidance
- Walking in the Implicit: Interactive World Exploration via Neural Scene RepresentationNeural Implicit Scene, VAE encoder, camera trajectories, diffusion transformer, geometry-aware retrieval, implicit state
- How Post-Training Shapes Biological Reasoning Modelscontinued pre-training, foundation models, generalization, in-domain performance, language models, multimodal biological data
- LISA: Likelihood Score Alignment for Visual-condition Controllable GenerationLISA, conditional control, decoder, diffusion models, disentangled features, feature projection
- MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image GenerationFID, Masked Image Modeling, Normalizing Flows, VAE encoder, generative flow, high-frequency synthesis
- Discretizing Reward ModelsMonte Carlo dropout, discretization, discriminative ability, oversensitivity, policy learning, reinforcement learning
- EO-WM: A Physically Informed World Model for Probabilistic Earth Observation ForecastingNDVI, Normalized Difference Vegetation Index, climatological baseline, cumulative physical stress signals, diffusion models, meteorological forcing
- Hallucination in World Models is Predictable and Preventablecoverage-aware sampling, curiosity rewards, data-centric signals, data-efficient fine-tuning, ground-truth actions, hallucination
- Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA EnhancementVision-Language-Action models, domain gap, imitation learning, object-centric representation, pose estimation, reinforcement learning
- The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of TatarTatar language, comparative experiments, cross lingual transfer, evaluation, fine tuning, low resource languages
- RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation6D Plucker ray coordinates, Vision Foundation Models, any-resolution cross-attention, dense prediction tasks, feature upsampling, geometry-aware neighborhood attention
- CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent EconomiesLLM agents, agent behavior, autonomous agents, communication, cumulative net income, economic systems
- Ko-WideSearch: A Korean Breadth-Search Benchmark for Exhaustive Set Enumeration by Web AgentsColumn-F1, Item-F1, Row-F1, automated synthesize-and-verify pipeline, breadth-search, composite key
- OpenBioRQ: Unsolved Biomedical Research Questions for Agentsagentic collapse, agentic models, answer key, biomedical research questions, citation verification, frontier agents
- Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoEMixture-of-Experts, compute allocation, denoising process, diffusion models, latent features, post-training framework
- Interleaved Speech Language Models Latently Work In Textintermediate layers, logit lens, speech language models, speech recognition, speech-text interleaving, spoken knowledge abilities
- NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement LearningMLLM-judged image quality, adjoint sensitivity analysis, classifier-free guidance, flow-based generators, forensic realism, hinge penalty
- The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction4D hand motion reconstruction, egocentric video, full frames, hand-overlay rendering, hand-pose annotations, metric-scale pose
- Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robotsattention masking, bi-manual manipulation, embodiment differences, head-camera frame, interleaved action tokens, parallel grippers
- Learning Transferable Dynamics Priors from Action to World Modelingaction-conditioned, diffusion world model, dynamics priors, multi-view interactive, policy-centric learning, pretraining
- PoseShield: Neural Collision Fields for Human Self-Collision ResolutionEikonal equation, SMPL, constrained optimization, motion generation, neural collision constraint, pose estimation
- Fast LeWorldModelJoint-Embedding Predictive Architectures, LeWorldModel, action-prefix prediction, autoregressive rollout, latent transition model, latent world model
- Confidence-Aware Tool Orchestration for Robust Video UnderstandingBlind Trust Problem, agentic video understanding, calibrated reliability score, confidence-cost GRPO reward, evidence interface, reliability-relevance score
- Geometric Stability of Neural Population Codes: Regional Variation, Behavioral Relevance, and Circuit DependenceSpearman rank correlation, attractor network model, feedforward input, geometric stability, hippocampal circuits, pairwise distance structure
- Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments3D reasoning, agent generalization, agentic systems, automated evaluation engine, benchmark, graphical understanding
- Beyond IID: How General Are Tabular Foundation Models, Really?Data Foundry, IID data, benchmarking, deep learning models, high-dimensional datasets, non-IID data
- Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty PredictionLarge Reasoning Models, cognitive episodes, difficulty prediction, effort allocation, episode sequences, episode-dynamic features
- Qwen-Image-2.0-RL Technical ReportGRPO-based RL training framework, chain-of-thought reasoning, hybrid classifier-free guidance, image editing, intra-group reward range filtering, on-policy distillation
- MultiHashFormer: Hash-based Generative Language ModelsHash Decoder, Hash Encoder, MultiHashFormer, Transformer decoder, causal LMs, discrete hash IDs
- PhysiFormer: Learning to Simulate Mechanics in World Space3D meshes, attention factorised, autoregressive baselines, denoising diffusion process, diffusion transformer, permutation-invariant
- In-Context World Modeling for Robotic ControlVision-Language-Action models, in-context adaptation, novel configurations, parameter updates, real-world robot platforms, robot policies
- ReasoningLens: Hierarchical Visualization and Diagnostic Auditing for Large Reasoning ModelsChain-of-Thought traces, agentic auditor, diagnostic auditing, hierarchical visualization, information necropsy, systemic reasoning profiles