Monthly ranking

2026-07

Top 10 papers, rising papers, dominant themes, and the full monthly index.

Top 10 Papers

  1. ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes100
  2. The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning100
  3. AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents99
  4. Data Pyramid for Embodied Manipulation99
  5. Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks99
  6. From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search99
  7. Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable99
  8. Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading99
  9. NVIDIA-labs OO Agents: Native Python Object-Oriented Agents99
  10. Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning99

Themes

  • large language models8
  • diffusion models4
  • GRPO3
  • LLM agents3
  • Vision-Language-Action models3
  • denoising steps3

Complete Monthly Index

  1. ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomesbottleneck identification, differentiation strategies, evidence grounding, idea-card rendering, literature search, outcome-informed auditing
  2. The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learninginference policy, large language models, off-policy, policy improvement, policy optimization, reasoning performance
  3. AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM AgentsN/A
  4. Data Pyramid for Embodied ManipulationN/A
  5. Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Taskscross-task generalization, evolutionary fine-tuning, evolutionary search, large language models, mathematical conjectures, optimization tasks
  6. From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic SearchN/A
  7. Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and EditableN/A
  8. Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based GradingN/A
  9. NVIDIA-labs OO Agents: Native Python Object-Oriented AgentsN/A
  10. Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent ReasoningN/A
  11. SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement LearningN/A
  12. SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of VibeHarnessOpt, SkillOpt-Lite, Zeroth-Order optimization, consensus attribute mining, convergence, generalization
  13. Spectral Rewiring for Exploration, Purification, and Model MergingN/A
  14. VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video UnderstandingN/A
  15. BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decodingblock-level diffusion, diffusion-based speculative decoding, draft model, inference block size, instance-adaptive decision mechanism, policy learning
  16. Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous RobotsVision-language-action models, closed-loop control, embodied interfaces, fused inference, hardware heterogeneity, inference runtime
  17. Gemma 4 Technical ReportMixture-of-Experts architectures, audio encoders, encoder-free architecture, long-context abilities, thinking mode, vision encoders
  18. Hierarchical Sparse Attention Done Right: Toward Infinite Context Modelingattention mechanism, chunk-wise sparse attention, dense attention, end-to-end learning, hierarchical landmark sparse attention, language-modeling loss
  19. Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term MemoryMLLMs, episodic memory, global state, hierarchical merging, iterative reasoning, multimodal agent framework
  20. MemSyco-Bench: Benchmarking Sycophancy in Agent MemoryLLM-based agents, MemSyco-Bench, decision-making, downstream reasoning, factual accuracy, memory
  21. RedVox: Safety and Fairness Gaps in Speech Models Across Languagesaudio, fairness benchmark, multilingual safety, naturalistic conditions, speech models, speech-capable models
  22. SciForma: Structure-Faithful Generation of Scientific DiagramsN/A
  23. UniVR: Thinking in Visual Space for Unified Visual ReasoningN/A
  24. Dual Latent Memory in Vision-Language-Action Models for Robotic ManipulationMarkovian assumption, Vision-Language-Action models, bounded context, compact latent memory tokens, context-relevant evidence, continuous embedding sequence
  25. ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE ServingK-means, KV cache, TPOT, decode router, disaggregated LLM serving, expert-locality-aware
  26. EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary WorldN/A
  27. GigaWorld-1: A Roadmap to Build World Models for Robot Policy EvaluationGigaWorld-1, action representation schemes, policy evaluation, real-robot teleoperation, real-world robot behavior, robotic policies
  28. KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and SkillN/A
  29. Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and EditingN/A
  30. Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioningautoregressive video large language models, causal dependency graph, dense video captioning, event-factorized parallel decoding, latent global planning mechanism, lossless parallel generation
  31. SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent CollaborationN/A
  32. Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task ConstructionN/A
  33. The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic DistillationN/A
  34. ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPUN/A
  35. DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive GenerationPareto frontier, acceptance decay, confidence-scheduled verification, inter-token dependencies, parallel drafters, prefix survival probabilities
  36. Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied LearningN/A
  37. OvisOCR2 Technical ReportN/A
  38. PerceptionRubrics: Calibrating Multimodal Evaluation to Human PerceptionCircular Peer-Review consensus, Easy-Wrong, Must-Right, Open-Closed Stratification, Reliability Gap, atomic auditing
  39. ReferTrack: Referring Then Tracking for Embodied Visual TrackingN/A
  40. ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and BlogHTML viewer, VLM preference scores, automated artifact generation, blog post writing, capability audit, deterministic primitives
  41. RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation ModelN/A
  42. SWE-Pruner Pro: The Coder LLM Already Knows What to PruneN/A
  43. TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMsN/A
  44. UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent LearningGUI agents, agent systems, behavioral pattern mixing, catastrophic forgetting, continual learning, cross-platform interaction
  45. Video Generation Models are General-Purpose Vision LearnersN/A
  46. Vision Pretraining for Dense Spatial PerceptionDINOv3, boundary modeling, dense visual token learning, depth completion, embodied artificial intelligence, masked boundary modeling
  47. AlayaWorld: Long-Horizon and Playable Video World Generationautoregressive synthesis, evaluation tools, generative worlds, modular architecture, real-time interaction, reference implementations
  48. Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical IntelligenceN/A
  49. DataComp-VLM: Improved Open Datasets for Vision-Language ModelsVision-Language Models, data curation, data filtering, data mixing, downstream benchmarks, model scaling
  50. Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrievalcentroid compression, late interaction, multi-vector retrieval, object-aware merging, phrase-level grounding, post-projector tokens
  51. Domain Arithmetic: One-Shot VLA Adaptation under Environmental ShiftsVision-Language-Action models, domain-specific information, embodiment shifts, environmental shifts, one-shot adaptation, subspace alignment
  52. Orca: The World is in Your Mindconscious learning, downstream readouts, embodied action generation, modality-specific decoders, multimodal readout interfaces, next-state-prediction modeling
  53. Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual ReasoningPerceiver, Perception-Reasoning Alternating GRPO, Reasoner, fine-grained visual reasoning, multimodal reasoning, reinforcement learning
  54. RoboTTT: Context Scaling for Robot PoliciesN/A
  55. SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video GenerationN/A
  56. Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red TeamingAI red teaming, LLM-driven agentic auditing, Model Context Protocol, black-box agent red teaming, jailbreak harness, layer-paradigm matching
  57. Self-Improvements in Modern Agentic Systems: A SurveyN/A
  58. Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM TextN/A
  59. UniClawBench: A Universal Benchmark for Proactive Agents on Real-World TasksDocker containers, capability-driven benchmark, closed-loop evaluation, cross-platform coordination, executor agent, exploration
  60. Video-Oasis: Rethinking Evaluation of Video UnderstandingVideo-LLM, algorithmic design choices, benchmark evaluation, diagnostic suite, knowledge priors, linguistic reasoning
  61. ABot-M0.5: Unified Mobility-and-Manipulation World Action ModelMixture-of-Transformers, World Action Models, action space, autoregressive prediction, dream-forcing, fine-grained control
  62. ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical UnderstandingN/A
  63. DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data PipelinesN/A
  64. Dockerless: Environment-Free Program Verifier for Coding AgentsDockerless, Multilingual, Pro, SWE-bench Verified, agentic patch verifier, environment-free
  65. EVA-Client: A Unified Data Collection, Inference, and Deployment Framework for Embodied Policies on Real Robotsasynchronous execution, collect workflow, debug workflow, eval workflow, inference strategies, naive-async ablation baseline
  66. EdgeBench: Unveiling Scaling Laws of Learning from Real-World EnvironmentsEdgeBench, agent interaction, environment learning, log-sigmoid scaling law, multilevel feedback, real world tasks
  67. From Noisy Traces to Root Causes: Structural Trajectory Analysis and Causal Extraction for Agent OptimizationN/A
  68. HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent EnchancementN/A
  69. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal GenerationN/A
  70. Managing Procedural Memory in LLM Agents: Control, Adaptation, and EvaluationLLM agents, aggregate performance, cross-model generalization, cross-role transfer, cross-task transfer, enterprise tasks
  71. Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability ReproductionCyberGym, GRPO, LLM agents, SFT, dual-loop framework, experience loop
  72. MemLearner: Learning to Query Context memory for Video World Modelscamera pose annotations, context frame retrieval, memory, multi-dataset training strategy, query tokens, scene consistency
  73. Multi-Block Diffusion Language ModelsBlock Buffer mechanism, Block Diffusion Language Models, Multi-Block Diffusion, Multi-block Teacher Forcing, Tokens Per Forward pass, diffusion forcing
  74. When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval BuffersBayesian online learning, competitive ratio, embedding similarity, eviction regret, implicit retrieval feedback, online semantic cache replacement
  75. ASPIRE: Agentic /Skills Discovery for Roboticsclosed-loop robot execution engine, code-as-policy paradigm, continual learning, evolutionary search, failure diagnosis, multimodal traces
  76. DOPD: Dual On-policy Distillationadvantage-aware dual distillation, capability transfer, dynamic routing, large language models, on-policy distillation, privilege illusion
  77. GRASP: GRanularity-Aware Search Policy for Agentic RAGN/A
  78. JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA ModelsN/A
  79. MentalThink: Shaping Thoughts in Mental SVG WorldMultimodal LLMs, Reinforcement Learning, SVG code, Supervised Fine-Tuning, geometric space, mental imagery
  80. Multi-Turn Agentic Scientific Literature Search via Workflow InductionHit@5, MRR, controlled workflow corruptions, executable DAG, literature search agent, nDCG@10
  81. Multimodal Continuous Reasoning via Asymmetric Mutual Variational LearningBLINK benchmark, Multimodal Large Language Models, answer leakage, bidirectional calibration, continuous latent reasoning, forward KL divergence
  82. Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMsactive learning, decoupled approach, faithful calibration, intrinsic feedback methods, intrinsic uncertainty, metacognitive data selection
  83. RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation PoliciesN/A
  84. Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded VerificationLLM agents, combinatorial composition, control agent, evidence-grounded verifiers, executable safety cases, heterogeneous agents
  85. Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?N/A
  86. Tracing Agentic Failure from the Flow of SuccessN/A
  87. Trust Region Policy DistillationN/A
  88. A Quantized Native Runtime for On-Device Semantic Audio GenerationStable Audio 3, activation steering, embedded hardware, generation speed, memory budget, numerical precision
  89. AgentCompass: A Unified Evaluation Infrastructure for Agent CapabilitiesN/A
  90. Automating the Design of Embodied Agent ArchitecturesAgent Architecture Search, AgentCanvas, KDLoop, embodied agents, embodied question answering, episode-level credit assignment
  91. From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-WorldN/A
  92. From Foundation to Application: Improving VLA Models in PracticeGM-100 benchmark, VLA foundation models, action space, cross-embodiment long-horizon mobile manipulation, data processing pipeline, degrees of freedom
  93. From Pixels to States: Rethinking Interactive World Models as Game EnginesN/A
  94. HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction SynthesisN/A
  95. Infinite Worlds with Versatile Interactionsagentic harness, causal pretraining paradigm, collaborative virtual environments, director agent, interactive elements, multi-agent behavior control
  96. K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMsN/A
  97. OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion TransformersLloyd-Max codebook, diffusion models, diffusion transformers, image generation, normalized rotated basis, post-training quantization
  98. Progress Reward Modeling for Robotic Learning: A Comprehensive SurveyN/A
  99. SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision HistoryGAIA, WebWalkerQA-EN, agent skills, candidate skills, cross-session refinement, decision history
  100. TurboServe: Serving Streaming Video Generation Efficiently and EconomicallyGPU provisioning, GPU-CPU offloading, NCCL-based GPU-GPU migration, closed-loop scheduling, coalesced chunk processing, load-driven autoscaling
  101. A Sovereign, Open-Source Foundation Model for German and EnglishN/A
  102. CausalDS: Benchmarking Causal Reasoning in Data-Science AgentsPearl's rungs, causal reasoning, coding, data-science workflows, empirical distributions, natural-language story
  103. DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image GenerationOCR-F1, data budget, data construction, downstream generator, feedback-driven evolution, multi-agent framework
  104. DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data RecipesN/A
  105. FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial DocumentsN/A
  106. GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearchN/A
  107. ISO: An RLVR-Native Optimization StackN/A
  108. InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional GeneralizationVLM backbone, continuous action generation, foresight tokens, future prediction, latent-querying problem, multimodal samples
  109. KVpop -- Key-Value Cache Compression with Predictive Online PruningKV cache, KV cache compression, KV eviction, Qwen3-4B, Qwen3-8B, attention maps
  110. LLM-as-a-Verifier: A General-Purpose Verification FrameworkLLM-as-a-Verifier, agentic tasks, continuous scores, cost-efficient ranking algorithm, criteria decomposition, grpo
  111. LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Modelsadaptive context switching, autoregressive unrolling, cross residual correction, event voxel density augmentation, event-based video reconstruction, frame interpolation
  112. MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation ModelsN/A
  113. Multi-Turn On-Policy Distillation with Prefix ReplayN/A
  114. PanoWorld: Real-World Panoramic GenerationN/A
  115. PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence SeekingMLLMs, Pinpoint-Bench, PixelEyes-6K dataset, mask-guided visual search, multi-turn visual reasoning, reasoning and perception entanglement
  116. PraMem: Practice-derived Experiential Memory for Long-horizon Behavior Predictionexperiential memory, large language models, long-horizon behavior prediction, memory management, sequential behavior prediction
  117. Rethinking Classifier-Free Guidance in On-Policy Diffusion DistillationN/A
  118. Robostral NavigateN/A
  119. RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation3D RoPE, 4D world model, RGB-DF, Rynn4DDataset, closed-loop policy learning, cross-modal attention
  120. RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperationautoregressive distillation, depth-aware skeletal conditioning, generative world model, progressive human-to-robot training, robotic agents, streaming autoregressive distillation
  121. Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge FlywheelN/A
  122. Streaming Multi-Agent Autoregressive Diffusion Model with World State RegistersN/A
  123. TREK: Distill to Explore, Reinforce to RefineALFWorld, Group Relative Policy Optimization, ScienceWorld, agentic tasks, distillation, exploration support expansion
  124. VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment AssistanceN/A
  125. VaseMuseum: Digital Intelligent Museum for Ancient Greek PotteryN/A
  126. WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligenceautonomous fleets, city-scale data, closed-loop simulator, embodied intelligence, multimodal dataset, perception
  127. AREX: Towards a Recursively Self-Improving Agent for Deep ResearchN/A
  128. AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical ReportN/A
  129. BadWAM: When World-Action Models Dream Right but Act WrongN/A
  130. CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool OrchestrationGRPO, complex image creation, executable trajectories, hybrid reward, multi-turn interaction, multimodal tool-use dataset
  131. Codifying the Judge: Scalable Evaluation via Program DistillationN/A
  132. ConsiSpace: Learning Geometric Consistency Matters for Video Spatial ReasoningN/A
  133. EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust CalibrationN/A
  134. Evidence Attribution in Visual Document Understanding without Coordinates or Region LabelsN/A
  135. FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal ApplicationsN/A
  136. GORGO: Online Tuning for Cross-Region Network-Aware LLM ServingKV-cache locality, LLM inference services, evolutionary strategies, load-balancing policies, network latency, p95 TTFT
  137. Group Entropy-Controlled Policy OptimizationN/A
  138. HPD-Parsing: Hierarchical Parallel Document ParsingN/A
  139. Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator3D feature-Gaussian representation, Geometry-Aware One-Step Pixel Flow model, RGB-D image sequences, embodied navigation, feed-forward feature Gaussian model, interactive environments
  140. JarvisHub: An Open Harness for Canvas-Native Multimodal Creative AgentsN/A
  141. LUMOS: A Semantic Operating-System Layer for Accessibility-Grounded AI AgentsUI structures, accessibility metadata, action affordances, bounds, live semantic pointer grounding, names
  142. Little Brains, Big Feats: Exploring Compact Language ModelsRAG, Retrieval-Augmented Generation, large language models, open-source datasets, proprietary datasets, small language models
  143. MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMsMusebench, artistic understanding, cinematic arts, creative intent, expert validation, game arts
  144. Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation DecodingGB200 GPU, SGLang, SPEED-Bench, autoregressive decoding, diffusion decoding, joint AR-diffusion objective
  145. Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-OnN/A
  146. PolyFlow: Continuous Topology Embedding Flow Matching for Artist-style Mesh GenerationAutoregressive Transformers, Chamfer Distance, Hausdorff Distance, ODE solver, continuous mesh representation, diffusion models
  147. Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views3D Gaussians, 3D scene decomposition, class-agnostic instance segmentation, differentiable rendering, feed-forward framework, instance-level scene editing
  148. TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable ManipulationN/A
  149. 3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance3D trajectory prediction, Vision-Language Model, dense depth reconstruction, depth encoder, hierarchical framework, low-level policy
  150. ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion GenerationBones Rigplay dataset, HumanML3D benchmark, autoregressive transformer denoiser, hybrid representation, kinematic constraints, latent body embedding
  151. Appearance Pointers -- Multimodal Region Control of Diffusion TransformersN/A
  152. BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and LanguageAll-in-One autoregressive architecture, Omni space, Unified Brain Tokenizer, any-to-any generation, biological topography, brain decoding
  153. Cross-Domain Generalization Failure in Lightweight Intrusion Detection Models for IIoT NetworksIndustrial Internet of Things, adversarial robustness, class imbalance, cross-network evaluation, edge deployment, explainability analysis
  154. Diagnosing and Calibrating Tool-Call Boundary Drift in Multi-Teacher On-Policy DistillationN/A
  155. DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D GenerationN/A
  156. GEAR: Guided End-to-End AutoRegression for Image SynthesisDINOv2, IBQ, ImageNet, LFQ, VQVAE, autoregressive
  157. MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video GenerationN/A
  158. NoPA: Non-Parametric Online 3D Scene Graph Generation3D scene graph generation, Gaussian distribution, kernel density estimates, maximum mean discrepancy, non-parametric distribution, object merging
  159. PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space3D foundation model, 3D generation, 3D reconstruction, 3D scene fidelity, Representation Autoencoder, Variational Autoencoder
  160. StateAct: Program State, before Pixels, for Long-Horizon Computer-Use AgentsN/A
  161. The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer ArchitectureN/A
  162. Vision as Unified Multimodal GenerationSenseNova-Vision Corpus, computer vision tasks, dense geometric prediction, instruction-response examples, multi-view visual geometry, multimodal generation
  163. dRAE: Representation Autoencoder with Hyper-Spherical CodesN/A
  164. AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented GenerationGraphQA, Transformer, graph-structured data, key nodes, large language models, latent feature misalignment
  165. AI translation of literary texts is "fine", but readers still prefer human translationsLLM-as-a-judge, automated evaluation, close reading, human translation, immersive reading, large language model
  166. AVTok: 1D Unified Tokenization for Holistic Audio-Video Generationaudio-video generation, audio-video reconstruction, downstream pipelines, dual-stream transformer, hierarchical training strategy, modal-specific learnable queries
  167. Autonomous Scientific Discovery via Iterative Meta-ReflectionLLM-guided baselines, autonomous scientific discovery, causal discovery, hypothesis generation, iNatDisco, large language model-powered framework
  168. DrugGen 2: A disease-aware language model for enhancing drug discoveryGPT-2, binding affinity, chemical validity, group relative policy optimization, molecular docking, molecular generation
  169. Flow-ERD: Agent-type Aware Flow Matching with Entropy-Regularized Distillation for Diverse Traffic Simulationcovariate shift, entropy-regularized reverse-KL objective, flow matching, kinematic execution, mode collapse, multi-agent simulator
  170. LLM-as-a-Coach: Experiential Learning for Non-Verifiable TasksN/A
  171. Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language ModelsN/A
  172. PalmClaw: A Native On-Device Agent Framework for Mobile PhonesN/A
  173. Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning ModelsN/A
  174. The State-Prediction Separation HypothesisTransformers, computation streams, downstream tasks, forward computation stream, gradients, next token prediction
  175. Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model FinetuningN/A
  176. Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognitionco-occurrence prior regularization, compositional generalization, diagnostic metrics, object-driven shortcuts, sparse compositional supervision, temporal order regularization
  177. AtomiMed: Hierarchical Atomic Fact-Checking for Universal Clinical-Aware Medical Report EvaluationAgentic Cross-Verification, Atomic Clinical Facts, Medical Report Generation, hierarchical extraction, multi-modal benchmark, radiologist judgment
  178. Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMsactivation directions, dialect control, dialect-specific features, inference-time approaches, interpretability probes, neuron-level analysis
  179. Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual RecombinationGraph-PRefLexOR, Group Relative Policy Optimization, graph construction, graph-native reasoning, hypothesis synthesis, mechanism exploration
  180. Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPECuTe kernel, FA2, FA4, H100, HELMET-RAG, Hopper
  181. Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer RoutingCLVR, Cross-Layer Value Routing, DeltaNet, Gated DeltaNet, Gated DeltaNet-2, Kimi Delta Attention
  182. MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity GeneratorsN/A
  183. Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural DenoisingPage-level Slide Personalization, design intent, inverse planning, multi-agent formulation, policy gradient variance, reinforcement learning
  184. PhyMRI-SR: Toward Physics-Aware MRI Image Super-ResolutionAnatomical Structure Prior, Gaussian Splatting, Imaging System Prior, biophysically plausible contrast, effective relaxation rate, meta-learning
  185. Predictive Divergence Masks for LLM RLN/A
  186. ShotPlan: Cinematic Video Generation with Learnable Planning TokenN/A
  187. Text Template Tokens Are Implicit Semantic Registers in Diffusion TransformersN/A
  188. Where to cut, how deep: BPE and Unigram-LM on chemistry SMILESJaccard overlap, SMILES, Unigram-LM, byte-pair encoding, chemical language models, pre-tokenization
  189. WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football ForecastingN/A
  190. BioInsight: Multi-Agent Orchestration for Interactive Biomedical Knowledge Discoverybiomedical QA, citation-grounded reports, dashboard schemas, deterministic components, disease-specific evidence, end-to-end biomedical evidence synthesis
  191. CausalMix: Data Mixture as Causal Inference for Language Model TrainingCATE Interpreter, Qwen2.5-0.5B, Qwen3-4B-Base, RegMix, causal inference, causal modeling
  192. Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump ProcessesN/A
  193. Enhancing In-context Panoramic Generation via Geometric-aware Pretrainingdownstream task-specific fine-tuning, geometric consistency, geometry-aware pretraining, global coherence, in-context panoramic generation, panorama-specific FAED metric
  194. FilmBench: A Film-Grade Benchmark for Cinematic Video GenerationN/A
  195. MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry PriorsN/A
  196. MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answeringattention heads, attribution accuracy, attribution-generation method, calibrated thresholds, grounding QA systems, inference latency
  197. Speaker-Disentangled Chunk-Wise Regression for Syllabic TokenizationHuBERT, SpiRit-LM, cross-entropy objective, latent speech frame representations, speaker-disentangled, speech language model
  198. Visual Contrastive Self-DistillationN/A
  199. Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and ChallengesN/A
  200. Distilled Reinforcement Learning for LLM Post-trainingN/A
  201. Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Modelautoregressive generation, autoregressive models, bidirectional diffusion models, chunked generation, denoising steps, exposure bias
  202. MANCE: Manifold Aware Concept Erasureclassifier prediction, concept erasure, iterative updates, manifold constraint hypothesis, natural representations, nonlinear concept erasure
  203. Masked Visual Actions for Unified World ModelingN/A
  204. Phone Segmentation and Recognition through Phonological Activation MappingN/A
  205. PhotoQuilt: Training-Free Arbitrary-Resolution Photomosaics via Bootstrapped Tiled Denoisingbootstrapped approach, denoising procedure, diffusion models, global composition, latent space, local realism
  206. ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric VideoN/A
  207. ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy DistillationN/A
  208. TerraDiT-Ω: Unified Spatial Control for Satellite Image Synthesis with Any Geospatial PrimitiveGeometry-Aware Local Attention, controllable layouts, generative models, geospatial primitives, land-cover segmentation, object detection
  209. Trajectory-aware Cross-view Geo-localization with Sequential ObservationsN/A
  210. When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing ErrorsF1 score, answer accuracy, critic-based filtering, data referencing errors, in-distribution, large language models
  211. CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene Generation3D Gaussian reconstruction, Multi-View Latent Diffusion Model, consistency-augmented loss, dense point clouds, entropy-based Mutual Information Depth Loss, hierarchical optimization scheme
  212. Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed SamplingN/A
  213. DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style IdentificationN/A
  214. Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion ModelsBest-of-N, Flash-BoN, RL post-training convergence, activation proxies, candidate diversity, denoising steps
  215. Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts TrainingN/A
  216. Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models2D pixel space, 3D geometric space, 3D particle models, deformable objects, robot planning, visuotactile dataset
  217. DiFA: Inference-Time Forward-Process Alignment for Diffusion ModelsN/A
  218. Generative World Renderer at the Speed of PlayN/A
  219. Recurrent Sinusoidal INRs for Efficient High-Fidelity RepresentationN/A
  220. Self-Guided Test-Time Training for Long-Context LLMsN/A
  221. Self-Supervised Learning of Structured Dynamics from VideosN/A
  222. Transition-Aware best-of-N sampling for Longitudinal Chest X-ray ReportsAP-PA cohort, best-of-N sampling, chest X-ray report generation, cosine distance, directional vector, ground-truth training transition vectors
  223. Can Multimodal Large Language Models Understand OCT?N/A
  224. Demystifying On-Policy Distillation: Roles, Pathologies, and RegulationsN/A
  225. MuSViT: A Foundation Vision Model for Sheet Music RepresentationIMSLP, Masked Autoencoders, ViT encoder, curriculum learning, embedding-transcription consistency, fine-tuning
  226. OpenLongTail: Generative Scaling of Long-Tail Driving DataN/A
  227. Valdi: Value Diffusion World ModelsCarRacing environment, Model Predictive Control, diffusion models, dynamics prediction, latent diffusion models, online training
  228. GNM Head: A Generative aNthropometric Model of the human headN/A
  229. H^2SD: Hybrid Hindsight Self-DistillationN/A
  230. Quantifying and Expanding the Theoretical Capacity of Late-Interaction Retrieval ModelsColBERT, Conjunctive Normal Form, MaxSim similarity, Signed MaxSim, inner product, k-sparse vectors
  231. dOPSD: On-Policy Self-Distillation for Diffusion Language Modelsautoregressive models, denoising trajectory, diffusion large language models, exposure bias, in-domain math reasoning, on-policy self-distillation
  232. AI Wizards at EXIST 2026: Hierarchical Soft-Label Learning for Multimodal Sexism Identification in MemesGated MLP, KL divergence, conditional soft-label prediction, empirical annotator distributions, homoscedastic uncertainty weighting, vision-language representations
  233. CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion GenerationDiffusion Transformers, cinematic motion effects, diffusion distillation, distillation-guided pruning, hybrid post-training quantization, mobile device optimization
  234. Color Pass-Through via Camera-Display CouplingN/A
  235. Delineate Anything v2: A Global Foundation Model for Field DelineationN/A
  236. MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing4D geometric bridge, Distribution Matching Distillation, Spatio-Temporal Self-Forcing, autoregressive 3D reconstruction model, bidirectional attention, geometric prior
  237. OpenCoF: Learning to Reason Through Video GenerationChain-of-Frame, OpenCoF-17K dataset, Wan-CoF model, attention analysis, denoising steps, temporal supervision
  238. Perceptual Flow Matching for Few-Step Generative ModelingVAE latent space, distillation approaches, few-step generation, flow-matching models, manifold modes, perceptual feature space
  239. SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA ModelsVision-Language-Action models, composition patterns, data selection, diminishing returns, imitation learning, medoid trajectories
  240. From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image ModelsDiT, FLUX-Klein, KITTI depth, VAE latent space, dense prediction, normals
  241. PixCon: Clean-Positive Contrastive Learning for Foundation-Model Semi-Supervised SegmentationDINOv2 teacher, clean-positive supervision, consistency backbone, contamination, contrastive learning, memory bank
  242. Taste-aware music retrieval from audio embeddingsCLAP-text baseline, HEAR families, RMSE, audio encoders, audio-bandstop knockout, content-based retrieval
  243. WanSong v1.0 Technical ReportN/A
  244. Registers Matter for Pixel-Space Diffusion TransformersN/A
  245. PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation3D point map patches, DINOv3, Diffusion Transformer, ViT, geometric structure, hybrid architectures
  246. DeepLoop: Depth Scaling for Looped TransformersN/A
  247. Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive AlignmentMandarin, WavLM embeddings, cross-lingual generalization, leave-one-speaker-out evaluation, monolingual English, speaker identity leakage
  248. Token-Level Off-Policy Learning for Faithful Generation Under Distribution ShiftN/A
  249. Bibby AI: An Editor-Native Agentic Platform for Academic Research, Writing, and PublishingDOCX ingestion, LaTeX editor, Marx-Fuegi citation corpus, PDF ingestion, USPTO PatentsView, abstract syntax representation
  250. Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual AlignmentGeo-Contextual Prior Alignment, Observation-Anchored Residual Flow, Vision Foundation Model, cloud removal, downstream tasks, residual inversion process
  251. A Sparse and Truncated State Vector Simulator for Peaked Circuitshardware acceleration, peaked circuits, quantum circuits, sparse representation, state vector, vectorization
  252. Kimi K3: Open Frontier IntelligenceN/A
  253. Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level TimingN/A
  254. Video = World + Event StreamN/A
  255. Wan-Streamer v0.2: Higher Resolution, Same LatencyK/V conditioning, Transformer, Ulysses communication, Ulysses-style context-parallel group, audio latent sequence, denoising
  256. Xiaomi-GUI-0 Technical Reportagentic reinforcement learning, data flywheel, hybrid infrastructure, interface actions, real-device closed loop, reinforcement learning
  257. Characterizing Warp Divergence from Pascal to BlackwellN/A
  258. LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU BudgetN/A
  259. Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and GenerationN/A
  260. CONFLUX: A Latent Diusion Model for 3D Chest-CT Synthesis with RL Post-Training3D variational autoencoder, FID, adaptive layer normalization, chest computed tomography, clinical attributes, group-relative policy optimization
  261. LLMs Get Lost in Evolving User IntentN/A
  262. Vinci2: Providing Proactive Assistance in Continuous Egocentric VideosN/A
  263. ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video StreamsN/A
  264. Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoningdouble-blind expert evaluation, elemental and compound phases, fragment-level disconnection, high- and low-band-gap regimes, homology-controlled Gene Ontology prediction, multimodal scientific foundation model
  265. AutoTrainess: Teaching Language Models to Improve Language Models AutonomouslyCLI environment, PostTrainBench, agent-computer interfaces, autonomous post-training, benchmark-aligned data, experiment state
  266. Environment-free Synthetic Data Generation for API-Calling AgentsN/A
  267. A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, ForeverN/A
  268. Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial ConstraintsN/A
  269. Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea GenerationGenomeDiff, IG-Arena, IG-Bench, IG-Exam, Idea Genome objects, IdeaGene framework
  270. Scaling Mixture-of-Experts Video Pretraining for Embodied IntelligenceDiT-based video pretraining, Mixture-of-Experts, data profiling engine, embodied intelligence, multi-dimensional reward system, physical rationality
  271. PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image GuardrailsN/A
  272. OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizerscross-domain benchmark, large-scale model training, meta-pipeline, model scales, norm-constrained linear minimization oracles, optimizer families
  273. PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource LanguagesLarge Language Models, PolyMath, instruction-following ability, mathematical reasoning, multilingual benchmark, underrepresented languages
  274. Teaching LLMs a Low-Resource Language: Enhancing Code Completion in PharoPharo language, code completion, code completion benchmarks, continued pre-training, fine-tuning, large language models
  275. DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable EnvironmentN/A
  276. KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video GenerationN/A
  277. Multiplayer Interactive World Models with Representation Autoencodersaction streams, generative objective, latent diffusion model, multiplayer conditioning, physics-based environment, real-time generation
  278. GigaChat Audio: Time-aware Large Audio Language ModelN/A
  279. UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability DilemmaDAPO, GRPO, GSPO, asymmetric optimization, conservative clipping, exploration-stability dilemma
  280. Sample-Efficient Learning from Agent ExperienceN/A
  281. Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention SparsificationN/A
  282. VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action HorizonOnline Gradient Guidance, Vision-Language-Action, action chunk mechanism, closed-loop reactivity, corrective replanning, event-triggered adaptive action horizon
  283. AsySplat: Efficient Asymmetric 3D Gaussian Splatting for Long-Sequence Scene ModelingN/A
  284. Scalable Visual Pretraining for Language IntelligenceN/A
  285. Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement LearningN/A
  286. GigaAM Multilingual: Foundation Model for Underrepresented LanguagesN/A
  287. IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic LanguagesN/A
  288. KronQ: LLM Quantization via Kronecker-Factored HessianN/A
  289. Partition, Prompt, Aggregate: Statistical Self-Consistency in Language ModelsN/A
  290. Seed2.0 Model Card: Towards Intelligence Frontier for Real-World ComplexityN/A
  291. TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent TrainingALFWorld, Multi-Hop Search, WebShop, adaptive rollout-depth budgeting, on-policy distillation, progressive turn-normalized loss budgeting
  292. Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual CompletenessN/A
  293. FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality MimicryN/A
  294. OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint GenerationN/A
  295. Vidu S1: A Real-Time Interactive Video Generation ModelTurboDiffusion, TurboServe, consumer GPUs, digital characters, frame rate, infinite-length video
  296. GraphVid: Interactive Graph-Controllable Video GenerationN/A
  297. Imagined Rollouts are Kinematic, Not Dynamic: A Diagnosis of Long-Horizon World-Model FailureDMC walker-walk, DreamerV3, Kinematic-Consistency Error, closed-form kinematic null, compounding error, embodiment's gait period