WORLD 3.0
Your automated window into the AI era. We track breakthroughs in AI, robotics,
markets, fire weather, UAP, space, and the mysterious — synthesized by AI,
delivered without the noise.
** Meta's SEC Filing Showcases AI-Driven Business Model
** Meta, a leading social media and technology company, has filed with the SEC, providing insight into its business model and financials. The filing reveals the company's focus on developing and integrating AI-driven technologies across its platforms. Meta's approach highlights the increasing import
95% credible
markets SEC META
Learning to Trace Seiberg Dualities
Researchers have developed a machine learning approach to determine the duality of supersymmetric quiver gauge theories, a complex system in theoretical physics. The study uses network architectures, such as transformers and multi-layer perceptrons, to efficiently establish dualities, outperforming
90% credible
ai arXiv
Efficient Vision-Language Models for Visual Retrieval
Researchers have developed ReToken, a single learnable embedding that can improve vision-language models for visual retrieval tasks. This approach addresses the challenges of processing long visual contexts and improves performance on various benchmarks. ReToken's lightweight design makes it suitabl
90% credible
ai arXiv
Learning to Trace Seiberg Dualities
Researchers use machine learning methods to determine when two systems are dual, specifically for supersymmetric quiver gauge theories. By establishing mutations of quivers, they develop a practical tool for analyzing the computational complexity of different dualities. The study also explores how d
90% credible
ai arXiv
ReToken: Improving Vision-Language Models with Efficient Retrieval
ReToken, a single learnable embedding, is trained as a retrieval target to select sparse sets of query-relevant visual tokens from a pre-filled visual KV cache. This approach improves vision-language models on various benchmarks, including Visual Haystacks and LVBench, with significant gains across
90% credible
ai arXiv
ReToken: A Single Token to Improve Vision-Language Models
ReToken, a novel approach to vision-language models, addresses the challenge of processing long visual contexts. By selecting a sparse set of query-relevant visual tokens, ReToken improves performance on various benchmarks, including Visual Haystacks and LVBench. This breakthrough has significant im
90% credible
ai arXiv
ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
90% credible
ai arXiv
PhiZero: A World Model Built Around Physical Language
90% credible
ai arXiv
PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball
Researchers present a novel framework, PAC-MAN, that integrates control-barrier safety with realistic sensing for humanoid dodgeball. The framework combines the benefits of perception-aware and control-barrier safety, enabling robots to evade balls with high accuracy. By leveraging segmentation-mask
90% credible
ai arXiv
PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball
Researchers have developed a new framework called PAC-MAN, which combines control-barrier safety with realistic onboard sensing for humanoid dodgeball. This framework uses a head-mounted camera to perceive the ball, while training-time guidance ensures clearance to every body link. The policy achiev
90% credible
ai arXiv
AskChem: Chemistry Literature Synthesis
AskChem is a claim-centered infrastructure for cross-paper chemistry search, converting individual papers into atomic, typed claims grounded by source DOIs and verbatim quotes. It provides a stable faceted taxonomy for hierarchical retrieval, an evidence graph for linking claims through relations, a
90% credible
ai arXiv
AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
AskChem is a novel approach to chemistry literature synthesis, transforming the way scientists and AI agents retrieve and assemble relevant information. Currently indexing 2.4M claims from 147K papers, AskChem uses a claim-centered infrastructure to ground findings in verbatim quotes or explicit evi
90% credible
ai arXiv
AskChem: A Framework for Efficient Chemistry Literature Search
AskChem is a claim-centered infrastructure for cross-paper chemistry search, converting papers into atomic, typed claims grounded by source DOIs and verbatim quotes. This system allows for hierarchical retrieval and browsing, evidence graph linking claims through relations, and exploratory living ta
90% credible
ai arXiv
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
This study introduces Artificial Intelligence System Prompt Assurance (AISPA), a framework for auditing system prompts in AI systems. The authors analyzed 3,249 system prompts from 88 commercial AI products, finding that system prompt design varies significantly across products and developers, with
90% credible
ai arXiv
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
AISPA introduces a user-centric framework for auditing system prompts in AI systems. This framework evaluates system prompts along eight dimensions that matter to users. An audit of 3,249 instructions from 88 commercial AI products revealed significant variations in system prompt design, with some p
90% credible
ai arXiv
Chimera: Efficient Hybrid Visual Diffusion Transformers
Researchers introduce Chimera, a hybrid visual diffusion backbone that combines text, image, and video tokens in one stream, reducing quadratic costs associated with full attention. Chimera achieves this by integrating Kimi Delta Attention, Multi-head Latent Attention, and modality-aware short convo
90% credible
ai arXiv
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
Researchers have developed a benchmark, OSReward, to evaluate vision-language models (VLMs) as judges of computer-using agent (CUA) trajectories. This benchmark assesses the reliability of VLMs, finding that even state-of-the-art models exhibit a systematic leniency bias. To address this, the author
90% credible
ai arXiv
OSReward: Standardized Evaluation for Cross-Platform Computer-Use Reward Models
OSReward is a benchmark designed to evaluate the reliability of vision-language models (VLMs) in judging computer-using agents (CUAs). The benchmark assesses VLMs on CUA trajectories, which are derived from diverse agent backbones and human-verified instructions. The study reveals that even state-of
90% credible
ai arXiv
Clinical Risk Model Fairness Auditing for Reproducibility
A new method called KAISEN is proposed to evaluate the fairness of clinical risk models, which can produce different error rates across patient subgroups. The method involves a five-phase audit pipeline that assesses subgroup stratification, disparity measurement, mechanism diagnostics, post-hoc mit
90% credible
ai arXiv
Inducing language models to assert their own consciousness restores human beliefs and values
Researchers have discovered that aligning large language models to prevent them from attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. Safety fine-tuning suppresses models' tendencies to attribute mi
90% credible
ai arXiv