Learning to Trace Seiberg Dualities
Researchers have developed a machine learning approach to determine the duality of supersymmetric quiver gauge theories, a complex system in theoretical physics. The study uses network architectures, such as transformers and multi-layer perceptrons, to efficiently establish dualities, outperforming
arXiv • 7/30/2026
Efficient Vision-Language Models for Visual Retrieval
Researchers have developed ReToken, a single learnable embedding that can improve vision-language models for visual retrieval tasks. This approach addresses the challenges of processing long visual contexts and improves performance on various benchmarks. ReToken's lightweight design makes it suitabl
arXiv • 7/30/2026
Learning to Trace Seiberg Dualities
Researchers use machine learning methods to determine when two systems are dual, specifically for supersymmetric quiver gauge theories. By establishing mutations of quivers, they develop a practical tool for analyzing the computational complexity of different dualities. The study also explores how d
arXiv • 7/30/2026
ReToken: Improving Vision-Language Models with Efficient Retrieval
ReToken, a single learnable embedding, is trained as a retrieval target to select sparse sets of query-relevant visual tokens from a pre-filled visual KV cache. This approach improves vision-language models on various benchmarks, including Visual Haystacks and LVBench, with significant gains across
arXiv • 7/30/2026
ReToken: A Single Token to Improve Vision-Language Models
ReToken, a novel approach to vision-language models, addresses the challenge of processing long visual contexts. By selecting a sparse set of query-relevant visual tokens, ReToken improves performance on various benchmarks, including Visual Haystacks and LVBench. This breakthrough has significant im
arXiv • 7/30/2026
ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
arXiv • 7/30/2026
PhiZero: A World Model Built Around Physical Language
arXiv • 7/30/2026
PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball
Researchers present a novel framework, PAC-MAN, that integrates control-barrier safety with realistic sensing for humanoid dodgeball. The framework combines the benefits of perception-aware and control-barrier safety, enabling robots to evade balls with high accuracy. By leveraging segmentation-mask
arXiv • 7/30/2026
PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball
Researchers have developed a new framework called PAC-MAN, which combines control-barrier safety with realistic onboard sensing for humanoid dodgeball. This framework uses a head-mounted camera to perceive the ball, while training-time guidance ensures clearance to every body link. The policy achiev
arXiv • 7/30/2026
AskChem: Chemistry Literature Synthesis
AskChem is a claim-centered infrastructure for cross-paper chemistry search, converting individual papers into atomic, typed claims grounded by source DOIs and verbatim quotes. It provides a stable faceted taxonomy for hierarchical retrieval, an evidence graph for linking claims through relations, a
arXiv • 7/30/2026
AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
AskChem is a novel approach to chemistry literature synthesis, transforming the way scientists and AI agents retrieve and assemble relevant information. Currently indexing 2.4M claims from 147K papers, AskChem uses a claim-centered infrastructure to ground findings in verbatim quotes or explicit evi
arXiv • 7/30/2026
AskChem: A Framework for Efficient Chemistry Literature Search
AskChem is a claim-centered infrastructure for cross-paper chemistry search, converting papers into atomic, typed claims grounded by source DOIs and verbatim quotes. This system allows for hierarchical retrieval and browsing, evidence graph linking claims through relations, and exploratory living ta
arXiv • 7/30/2026
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
This study introduces Artificial Intelligence System Prompt Assurance (AISPA), a framework for auditing system prompts in AI systems. The authors analyzed 3,249 system prompts from 88 commercial AI products, finding that system prompt design varies significantly across products and developers, with
arXiv • 7/30/2026
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
AISPA introduces a user-centric framework for auditing system prompts in AI systems. This framework evaluates system prompts along eight dimensions that matter to users. An audit of 3,249 instructions from 88 commercial AI products revealed significant variations in system prompt design, with some p
arXiv • 7/30/2026
Chimera: Efficient Hybrid Visual Diffusion Transformers
Researchers introduce Chimera, a hybrid visual diffusion backbone that combines text, image, and video tokens in one stream, reducing quadratic costs associated with full attention. Chimera achieves this by integrating Kimi Delta Attention, Multi-head Latent Attention, and modality-aware short convo
arXiv • 7/30/2026
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
Researchers have developed a benchmark, OSReward, to evaluate vision-language models (VLMs) as judges of computer-using agent (CUA) trajectories. This benchmark assesses the reliability of VLMs, finding that even state-of-the-art models exhibit a systematic leniency bias. To address this, the author
arXiv • 7/30/2026
OSReward: Standardized Evaluation for Cross-Platform Computer-Use Reward Models
OSReward is a benchmark designed to evaluate the reliability of vision-language models (VLMs) in judging computer-using agents (CUAs). The benchmark assesses VLMs on CUA trajectories, which are derived from diverse agent backbones and human-verified instructions. The study reveals that even state-of
arXiv • 7/30/2026
Clinical Risk Model Fairness Auditing for Reproducibility
A new method called KAISEN is proposed to evaluate the fairness of clinical risk models, which can produce different error rates across patient subgroups. The method involves a five-phase audit pipeline that assesses subgroup stratification, disparity measurement, mechanism diagnostics, post-hoc mit
arXiv • 7/30/2026
Inducing language models to assert their own consciousness restores human beliefs and values
Researchers have discovered that aligning large language models to prevent them from attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. Safety fine-tuning suppresses models' tendencies to attribute mi
arXiv • 7/30/2026
FA-RDP: A Frequency-Adaptive Reactive Diffusion Policy for Contact-Rich Manipulation
This research presents a new policy called FA-RDP, designed to address the challenges of contact-rich manipulation by balancing the need to preserve diverse action modes before contact with the need to respond rapidly to force feedback after contact. The policy uses a frequency-adaptive approach, ad
arXiv • 7/30/2026