Entities · Locations
arXivLabs
471 articles tagged with this entity.
-
Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety
-
DiffCVE: Diffusion-based Compressed Video Enhancement
-
An optimal control approach for neural network architecture adaptation with a posteriori error estimation
-
The Appeal and Reality of Recycling LoRAs with Adaptive Merging
-
Face-trace: Open-Set Attribution and Progressive Discovery of Synthetic Face Generators
-
Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents
-
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
-
AirPASS: Over-the-Air Federated Learning via Pinching Antenna Systems
-
CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views
-
Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics
-
Optimal Conformal Prediction under Epistemic Uncertainty
-
What Predicts Correctness in Text-to-SQL? A Selective-Prediction Study
-
Generalist Vision-Language Models for Fast Radio Burst detection: a zero-shot benchmark against a specialized detector
-
Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild
-
Quantifying Frontier LLM Capabilities for Container Sandbox Escape
-
Faithful or Findable? Evaluating LLM-Generated Metadata for RDF Dataset Search
-
BabyVision: Visual Reasoning Beyond Language
-
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
-
MoWorld: A Flash World Model
-
Beyond Refusal: A Same-Lineage Study of Aligned and Abliterated LLMs for Vulnerability Analysis
-
FourTune: Towards Fully 4-Bit Efficient Post-Training for Diffusion Models
-
Channel-wise Retrieval for Multivariate Time Series Forecasting
-
Harrison.Rad 1.5 Technical Report: A radiology foundation model that can draft reports from images, priors and clinical context
-
Information Limits and Attractor Dynamics in Economies of Frontier LLM Agents: A Pre-Registered Test
-
From Textural Counterpoint to Feature Encoding: A Multi-Dimensional Machine Representation Study of Haydn's "The Lark" Integrating Electroacoustic Analysis
-
Mitigating Errors in LLM-Generated Web API Invocations via Retrieval-Augmented Generation and Constrained Decoding
-
Spider 2.0-AIFunc: Extending Real-World Text-to-SQL to AI-Native SQL Workflows
-
Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation
-
From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation
-
Agentic Artificial Intelligence for Multistage Physics Experiments at a Large-Scale User Facility Particle Accelerator
-
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
-
ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents
-
A Retrieval-Augmented Framework for Detecting and Resolving Pragmatic Ambiguities in Natural Language Requirements
-
AutoCedar: An Agentic Framework for Verifier-Guided Access Control Policy Synthesis
-
Don't Wait to Reply: Towards Responsive yet Thoughtful Dialogue through Proactive Thinking
-
Predicting the Emergence of Induction Heads in Language Model Pretraining
-
Language Models as Higher-Order Planning Formalizers
-
TAMA: A Human-AI Collaborative Thematic Analysis Framework Using Multi-Agent LLMs for Clinical Interviews
-
Geometry of Ordinal Representations in Language Models
-
Quantize the Target, Quantize the Drafter: Efficient Inference with Qwen3.5-4B
-
Differentiate the Evaluator, Not the Program: An Efficient Runtime Representation for Neuro-Symbolic Learning
-
Rethinking AI-Generated Text Detection: A Strong Baseline and the Distribution-Shift Problem That Remains
-
Knowledge-Centric Information Systems
-
When Does Small Data Work? Accuracy and Efficiency Trade-offs Between Tabular Foundation Models and Conventional Methods for Crowd-State Classification at Hajj and Umrah
-
SMOCS: A Streaming Framework for Simplified Deployment, Monitoring, and Optimization of ML Systems in Production
-
Latent Programming Horizons in Coding Agents
-
Learning Flexible Generalization in Video Quality Assessment by Bringing Device and Viewing Condition Distributions
-
Repurposing CLIP to Localize at Pixel Level
-
Towards Realistic Remote Sensing Dataset Distillation with Discriminative Prototype-guided Diffusion
-
SIMPLER: Efficient Foundation Model Adaptation via Similarity-Guided Layer Pruning for Earth Observation
-
SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning
-
CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving
-
In-span learning: adapting reduced-order models using their own predictions
-
SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
-
StarTSE: Towards Streaming Target Speaker Extraction via Chunk-wise Interleaved Splicing of Autoregressive Language Model
-
SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
-
RAF: Reliability-Aware Fusion of Camera, LiDAR, and 4D RADAR for Robust 3D Object Detection in Adverse Weather
-
ProbeLogits: Kernel-Level LLM Inference Primitives for AI-Native Operating Systems
-
GLM-5 Serving Parameter Tuning for OpenClaw: Single-Deployment MaaS Inference Optimization for Long-Context Agent Workloads
-
Metronome: Bound the Cache, Keep the Beat for Real-Time Interaction Model Serving
-
Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models
-
Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI
-
Flow-A11y: Flow-Aware Accessibility Testing
-
Builder, Defender, Breaker: The Case Against Removing the Human from the AI-Driven Security Lifecycle
-
Semantic Segmentation-Driven Image-Level Diagnosis of Liver Cancers in Hematoxylin and Eosin Histopathology Images
-
Scalable Maximal Frequent Episode Mining with Desbordante
-
The Map Behind the Flow: Finite-Step Gradient Descent as a Dynamical System
-
Last-Meter Precision Navigation for UAVs: A Diffusion-Refined Aerial Visual Servoing Approach
-
From Open Loop to Closed Loop: A Test-Time Iterative Optimization Framework for Reference-Consistent Image Generation
-
An Exploration of Agentic Information Fusion for Test Maintenance Prediction
-
Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
-
Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure
-
Seduced by the Narrative: Assessing Rule Adherence in Semi-Open Textual Sandboxes
-
Attention Limited Reward Learning
-
Lacuna Inc. at SemEval-2026 Task 4: Structurally Gated State-Space Models for Disentangling Narrative Similarity
-
MIRAGE: Defending Long-Form RAG Against Misinformation Pollution
-
When Do Foundation Models Pay Off? A Break-Even Analysis of Pretrained Time Series Forecasters
-
Human-Centric Reflective Architecture for Human-AI Collaborative Decision-Making
-
Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5
-
Dashboard2Code: Evaluating Multimodal Models on Reconstructing Interactive Dashboards
-
Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
-
Benchmarking API Drift in LLM-Generated Quantum Code Across Successive SDK Versions
-
AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression
-
Reward Granularity in RLVR: Comparing Process and Outcome Reward Structures for Mathematical Reasoning in Small Language Models
-
GenHOI: Generalized Hand-Object Pose Estimation with Occlusion Awareness
-
Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment
-
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
-
Janus: a Playground for User-Involved Agentic Permission Management
-
Epistemic Goggles: A Pretrained Module that Induces an Epistemic Frame via Gradient Editing
-
Distributed Attacks in Persistent-State AI Control
-
PPTArena: A Benchmark for PowerPoint Editing
-
A global predicted-fMRI drive signal from TRIBE does not predict YouTube replay heatmaps
-
Split-n-Chain: Privacy-Preserving Multi-Node Split Learning with Blockchain-Based Auditability
-
AdaCount: Training-Free Similarity-Guided Spatial and Feature Adaptation for Zero-Shot Object Counting
-
A global optimization SAR image segmentation model can be easily transformed to a general ROF denoising model
-
ExPerT: Personalizing LLM Responses to Users' Domain Expertise via Query-Wise Semantic and Keystroke Behavioral Cues
-
Grounded Optimization: A Layered Engineering Framework for Reducing LLM Hallucination in Automated Personal Document Rewriting
-
Has This Checkpoint Been Abliterated? A Two-Signal Audit and Its Failure Map
-
BRIDGE: Predicting Human Task Completion Time From Model Performance
-
Beyond Skepticism: Evaluating LLMs Pedagogical Intent Reasoning with the Adaptive Pedagogical Vigilance Framework
-
Multi-THuMBS: Multi-person Tracking of 3D Human Meshes Beyond Video Shots
-
COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows
-
Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification
-
EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments
-
The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits
-
PreScience: A Dataset and Benchmark for Scientific Forecasting
-
Gaming Consensus: Coordinated Manipulation in Crowdsourced Fact-Checking
-
What Types of Human-AI Teams Exist?
-
GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation
-
SemEval-2026 Task 9: Detecting Multilingual, Multicultural and Multievent Online Polarization
-
GPTKB v1.5: A Massive Knowledge Base for Exploring Factual LLM Knowledge
-
Learning dynamical systems from noisy data with Weak-form Kernel Ridge Regression
-
Hardening x402: PII-Safe Agentic Payments via Pre-Execution Metadata Filtering
-
Planning over MAPF Agent Dependencies via Multi-Dependency PIBT
-
From Signals to Structure: How Memory Architecture Drives Language Emergence in LLM Agents
-
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity
-
Quantifying the Affective Gap: A Zero-Shot Evaluation of LLMs on Fine-Grained Emotion Taxonomies
-
From "Strings" to "Things" for Personal Knowledge Graphs: Evaluating LLM Triple Extraction for Recommendation Systems
-
AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation
-
NeuroCogMap Reveals Cognitive Organization of Large Language Models
-
LLVM-Bench: Benchmarking and Advancing Large Language Models for LLVM Compiler Issue Resolution
-
Crystalite: A Lightweight Transformer for Efficient Crystal Modeling
-
Explaining Tabular Foundation Model Differences Through Meta-Features
-
Comparing Large Language Models on Scrum Certification-Style Questions: Accuracy, Stability, and Error Patterns
-
A Text-Steerable Instrument for Sketching Procedural Soundscapes via Language Models
-
Prompting GPT-5 on Scrum Certification Questions: An Empirical Accuracy Study
-
Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences
-
Position: Collaborative Agentic AI Needs Interoperability Across Ecosystems
-
The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes
-
Large Databases Need Small, Open-Weight Language Models
-
Revealing Safety-Critical Scenarios for UTM via Transformer
-
Bridging Local Observation and Global Simulation in Closed-Loop Traffic Modeling
-
CVE-TTP KG: Knowledge Graph Linking Software Vulnerabilities to Attack Behaviors
-
Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles
-
Efficient Public Verification of Private ML via Regularization
-
Why Solve It Twice? Hierarchical Accumulation of Skills for Transfer-Efficient ML Engineering
-
Temporal Preservation over Processing: Diagnosing and Designing Spatiotemporal Single-Stage Video Detectors
-
FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents
-
Skin-R1: Clinical Knowledge-Guided Dermatological Diagnosis Using Vision-Language Models
-
StackingNet: Collective Inference Across Independent AI Foundation Models
-
Agentic Tool Use in Large Language Models
-
Compressed Sensing for Capability Localization in Large Language Models
-
Momentum Guidance: Plug-and-Play Guidance for Flow Models
-
Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction
-
Is Lying an Emergent Behaviour in LLMs? Evidence from Gaslighting AI agents in a Sustainability Game
-
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks
-
LLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive Dashboard
-
BackTranslation2.0 -- A Linguistically Motivated Metric to Assess Sign Language Production
-
Geometric Measurements of the Axiom of Choice in Neural Proof Embeddings
-
On the Nonlinearity of Learning Rate Scaling for LLM Training
-
Aristotelian Virtue Profiling of LLMs through Ethical Dilemmas
-
Entity Binding Failures in Tool-Augmented Agents
-
Few-class Fidelity: Evaluating Explanations of Real-conditions CNN classifiers with Optimized Perturbations
-
Trust Your Instincts: Confidence-Driven Test-Time RL for Vision-Language-Action Models
-
ASTAD: Asymmetric Style Transfer for Synthetic-to-Real Adaptation in Autonomous Driving
-
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing
-
UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models
-
Seed-to-Seed: Unpaired Image Translation in Diffusion Seed Space
-
UltraImageGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention
-
Beyond the Reranker: Do RAG Retrieval Enhancements Help Once a Strong Reranker Is Present?
-
Unlocking the Visual Record of Materials Science: A Large-Scale Multimodal Dataset from Scientific Literature
-
Experience Graphs: The Data Foundation for Self-Improving Agents
-
MCP Server Architecture Patterns for LLM-Integrated Applications
-
Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization
-
Arko-T: A Foundation Model for Text-to-Structured 3D Generation
-
Towards Human-Level Book-Writing Capability
-
EfficientUICoder: A Bidirectional Token Compression Framework for Efficient MLLM-Based UI Code Generation
-
The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction
-
Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving
-
How Far Can You Get Without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation
-
DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information
-
Reported Confidence in LLMs Tracks Commitment More Than Correctness
-
Foundation vs. Specialized Models: Evaluating Catastrophic Forgetting in Continual Time Series Forecasting
-
Neural Image Space Tessellation effect
-
HunyuanImage 3.0 Technical Report
-
Unified Zero-Shot Time Series Forecasting: A Darts Foundation
-
Towards Automating Scientific Review with Google's Paper Assistant Tool
-
When Is an LLM Worth It for Hyperparameter Optimization? A Budget-Matched Study on Tabular Data Finds the Warm-Start Is a Default Configuration, Not the Model
-
On the Effect of Uncertainty on Layer-wise Inference Dynamics
-
NormAct: A Benchmark for Hidden Social Norm Compliance in Embodied Planning
-
Radar Guided Camera Verification for Automatic Emergency Braking Rethinking Object Detection in Radar Camera Fusion
-
NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning
-
Seven Security Challenges That Must be Solved in Cross-domain Multi-agent LLM Systems
-
ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents
-
Symmetry-Aware Transformer Training for Automated Planning
-
ZooClaw-FashionSigLIP2: Distilled Fine-tuning for Robust Fashion Retrieval
-
Scene and Human in One World: Reconstruction in a Feedforward Pass
-
RSD: Moving Local Triangular Charts for Auditing Language-Model Hidden States
-
Contagion Networks: Evaluator Preference Propagation in Multi-Agent LLM Systems
-
SHIFT: Gate-Modulated Activation Steering for Knowledge Conflict Mitigation in Retrieval-Augmented Generation
-
Qwen-Image-2.0-RL Technical Report
-
ToE: A Hierarchical and Explainable Claim Verification Framework with Dynamic Multi-source Evidence Retrieval and Aggregation
-
Deployment-Side Adaptiveness in Multi-Horizon Volatility Forecasting
-
Does Aurora Encode Atmospheric Structure? Latent Regime Analysis and Attribution
-
RecallRisk-BERT: A Multi-Task Framework for Post-Report Medical Device Recall Triage
-
A3C3: AI Algorithm and Accelerator Co-design, Co-search, and Co-generation
-
Target-Aware Bandit Allocation for Scalable Surrogate Optimization in Chemical Space
-
How Good Can Linear Models Be for Time-Series Forecasting?
-
Hybrid privacy-aware semantic search: SVD-truncated document geometry and CKKS-encrypted query reranking under a restricted threat model
-
Learning State-Tracking from Code Using Linear RNNs
-
PMDformer: Patch-Mean Decoupling Information Transformer for Long-term Forecasting
-
Weak-to-Strong Elicitation via Mismatched Wrong Drafts
-
Do Image Editing Models Understand Lighting?
-
Boundary-Aware Context Grounding for A Low-Channel EEG Agent
-
Linguistics and Human Brain: A Perspective of Computational Neuroscience
-
Einstein World Models
-
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
-
CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs
-
OmniRobotHome: A Multi-Camera Home Platform for Real-Time Human-Robot Interaction
-
Use What You Know: Causal Foundation Models with Partial Graphs
-
Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets
-
PhoneBuddy: Training Open Models for Agentic Phone Use
-
Erased, but Not Gone: Output Forgetting Is Not True Forgetting
-
Scalable Peptide Design via Memory-Efficient Equivariant Transformer
-
Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos
-
Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment
-
ACT-JEPA: Novel Joint-Embedding Predictive Architecture for Efficient Policy Representation Learning
-
Robustness assessment of large audio language models in multiple-choice evaluation
-
Flexible Gravitational-Wave Parameter Estimation with Transformers
-
Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs
-
RoboAtlas: Contextual Active SLAM
-
VENI: Variational Encoder for Natural Illumination
-
A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language Models
-
Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding
-
GeoRanker: Distance-Aware Ranking for Worldwide Image Geolocalization
-
AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning
-
Latent Visual States for Efficient Multimodal Reasoning
-
MyoInteract: A Framework for Fast Prototyping of Biomechanical HCI Tasks using Reinforcement Learning
-
Average Rankings Mask Per-Subject Optimality: A Friedman-Nemenyi Benchmark of EEG Motor-Imagery BCI Decoders
-
Pigeonholing: Bad prompts hurt models to collapse and make mistakes
-
Paying to Know: Micro-Transaction Markets for Verified Product Information in Agentic E-Commerce
-
BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents
-
An adaptive subsampling method for large-sample feature screening
-
Which Spaces can be Embedded in $L_p$-type Reproducing Kernel Banach Space? A Characterization via Metric Entropy
-
Impatient Bandits: Optimizing for the Long-Term Without Delay
-
You Don't Need to Run Every Eval
-
HyMaTE: A Hybrid Mamba and Transformer Model for EHR Representation Learning
-
ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMs
-
Diffusion Integrated Gradients: Controllable Path Generation for Flexible Feature Attribution
-
RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems
-
VeriPilot: An LLM-Powered Verilog Debugging Framework
-
AdversaBench: Automated LLM Red-Teaming with Multi-Judge Confirmation and Cross-Model Transferability
-
The Measurable Majority
-
BioPIE: A Biomedical Protocol Information Extraction Dataset for Experiment Understanding
-
MassSpecGym in the Wild: Uncovering and Correcting Evaluation Pitfalls in AI-Driven Molecule Discovery
-
Holo-World: Unified Camera, Object and Weather Control for Video World Model
-
A High-Resolution Landscape Dataset for Concept-Based XAI With Application to Species Distribution Models
-
Self-Preference Is Weak or Absent in Verifiable Instruction-Following Revision: A Four-Model Test Under Genuine Authorship
-
A Multi-Agent system for Multi-Objective constrained optimization
-
Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems
-
Flow Matching for Efficient and Scalable Data Assimilation
-
HGCN(O): A Self-Tuning GCN HyperModel Toolkit for Outcome Prediction in Event-Sequence Data
-
REDACT: A Systematically Controlled Multilingual Benchmark for Personal Information Detection
-
Pseudo-Formalization for Automatic Proof Verification
-
Clusters are All You Need: Pre-Training the Tsetlin Machine with Semantic Clusters from Language Models for Interpretability
-
Stitching and dimensionality effects on large artificially generated volume datasets
-
Shape of Thought: Progressive Object Assembly via Visual Chain-of-Thought
-
Bidirectional Tutoring for Developmental Motor Learning in Robots: Co-Developed Interaction Dynamics Support Stable Learning
-
Simulation of Language Evolution under Regulated Social Media Platforms: A Synergistic Approach of Large Language Models and Genetic Algorithms
-
Charting the Future of Scholarly Knowledge with AI: A Community Perspective
-
Self-Adaptive Scale Handling for Forecasting Time Series with Scale Heterogeneity
-
Training-Free Metrics for Synthetic Object Detection Data: A Proxy for Detector Performance
-
Can Agents Distinguish Visually Hard-to-Separate Diseases in a Zero-Shot Setting? A Pilot Study
-
AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QA
-
MetaResearcher: Scaling Deep Research via Self-Reflective Reinforcement Learning in Adversarial Virtual Environments
-
Advancing DialNav through Automatic Embodied Dialog Augmentation
-
Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts
-
Towards Scalable Customization and Deployment of Multi-Agent Systems for Enterprise Applications
-
Hallucination Detection and Correction in Medical VLMs via Counter-Evidence Verification
-
HeatKV: Head-tuned KV-cache Compression for Visual Autoregressive Modeling
-
JourneyFormer: Encoding Airbnb Guest Journey with Sequence Modeling
-
CEO-Bench: Can Agents Play the Long Game?
-
As Easy as Rocket Science: Assessing the Ability of Large Language Models to Interpret Negation in Figurative Language
-
RouteJudge: An Open Platform for Reproducible and Preference-Aware LLM Routing
-
Dynamic In-Group Persona Generation for Enhancing Human-AI Rapport
-
G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom Alignment
-
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
-
Sensor Configuration Matters: A Systematic Evaluation of Multimodal SLAM on Quadruped Robots
-
Physics-IQ Verified
-
OpenAnt: LLM-Powered Vulnerability Discovery Through Code Decomposition, Adversarial Verification, and Dynamic Testing
-
Scalable Batch Bayesian Optimization Via Subspace Acquisition Functions
-
Riemannian MeanFlow for One-Step Generation on Manifolds
-
Fluently Lying: Adversarial Robustness Can Be Substrate-Dependent
-
EgoCS-400K: An Egocentric Gameplay Dataset for World Models
-
Attention Alignment Between Humans and Vision-Language Models
-
Surveying GenAI-based Automation in Printed Circuit Board Design and Test
-
Comprehensive pKa Data Augmentation from Limited Real Data through an Engineered Models-Quantum Framework
-
Implicit vs. Explicit Prompting Strategies for LVLMs in Referential Communication
-
Reward hacking in physical reinforcement learning revealed by turbulent drag reduction
-
Handling Feature Heterogeneity with Learnable Graph Patches
-
When LLMs Analyze Scars: From Images to Clinically-Meaningful Features
-
Planning with the Views
-
Complex Layout Classification in the Wild: A Low-Resource Approach with Layout-Preserving Augmentations
-
Seeing Is Not Screening: Multimodal Hidden Instruction Attacks on Agent Skill Scanners
-
Advances in 4D Representation: Geometry, Motion, and Interaction
-
A Benchmark for Omni-Modal Reasoning in Long Videos
-
GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
-
Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond
-
Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
-
Toward Accessible Psychotherapy Training Using AI-Driven Interactive Patient Avatars
-
EngTrace: A Symbolic Benchmark for Verifiable Process Supervision of Engineering Reasoning
-
Disentangling Perception and Reasoning in Multimodal LLMs via Reward Design
-
Ensemble RL through Classifier Models: Enhancing Risk-Return Trade-offs in Trading Strategies
-
PrologMCP: A Standardized Prolog Tool Interface for LLM Agents
-
Attribute Inference from Interactive Targeted Ads
-
Generative Molecular Design with Steerable and Granular Synthesizability Control
-
MacrOData: New Benchmarks of Thousands of Datasets for Tabular Outlier Detection
-
Unifying Post-hoc Explanations of Knowledge Graph Completions
-
The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling
-
Improved Knowledge Distillation for Land-Use Image Classification
-
HemExp: Clinically-Guided Latent Diffusion for Modeling Hematoma Expansion
-
Domain-Guided Prompting of the Segment Anything Model for Seismic Interpretation: The Role of Attributes, Visualization, and Hybrid Prompts
-
VinQA: Visual Elements Interleaved Long-form Answer Generation for Real-World Multimodal Document QA
-
Structure-aware Knowledge-guided Heterogeneous Mamba for Zygomaticomaxillary Suture Assessment
-
PROSE: Training-Free Egocentric Scene Registration with Vision-Language Models
-
Token-Level Entropy Reveals Demographic Disparities in Language Models
-
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation
-
Communication-Efficient Verifiable Attention for LLM Inference
-
MUNI: Multimodal Unified Latent Diffusion for Coherent Any-to-Any Generation
-
GPT-Based Fast Simulation of CLAS12 Detector Hits via Conditional Autoregressive Generation
-
STAR-NT: Spatiotemporal Acceleration of Real-Time Neural Transparency Rendering
-
HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
-
Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw
-
Surpassing Scale by Efficiency: A Compact 135M Parameter Foundational LLM Natively Adapted for the Bangla Language
-
Not What, But How: A Framework for Auditing LLM Responses across Positioning, Generalization, Anthropomorphism, and Maxims
-
Double-Helix Vision (DH-V2): A Geometry-Based Visual Sampler for Bandwidth-Constrained Perception
-
MatchLM2Lite: A Scalable MLLM-to-Lite Framework for Reproduced Content Identification
-
Is Code Better Than Language for Algorithmic Reasoning
-
MapDream: Task-Driven Map Learning for Vision-Language Navigation
-
False Sense of Safety in Selective Signal Classification: Auditing Bound Tightness and Exchangeability for Risk Control
-
Graphical-Probabilistic Modeling of Generative Flows in LLM-Native Software Systems
-
AI Supply Chain Galaxy: 3D Visual Analytics for License Compliance
-
MultiMolecule: a modular ecosystem for biomolecular sequence-model workflows
-
LiFT: Local Search via Linear Programming for Overfitting-Controlled Transformers
-
The Answer Lies Within: Self-Derived Rewards Enable Explainable Relation Extraction
-
Protein Design with Agent Rosetta: A Case Study for Specialized Scientific Agents
-
Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework
-
Best Arm Identification with Minimal Regret
-
FraudSMSWalker: Benchmarking Agentic Large Language Models for SMS-to-Webpage Fraud Detection
-
RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments
-
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
-
Human genetic evidence is associated with drug approval across therapeutic areas: an observational analysis of 26,278 target-disease pairs with temporal validation and feature ablation
-
AutoDojo: Adaptive Attacks Expose Superficial Defenses and User-Underspecification Limits in LLM Agents
-
Simplifying the Modeling of Arbitrary Conditionals in Natural Language
-
When Do We Need LLMs? A Diagnostic for Language-Driven Bandits
-
A Pragmatic VLA Foundation Model
-
A Robust Point Cloud Analysis Framework Inspired By Primary Visual Cortex
-
S$^2$COPE: Self-Supervised Concept Discovery via Preference Learning
-
Crypto x AI, AI x Crypto: A Survey
-
A Statistical and Machine Learning Framework for Operational Threshold Detection and Deployable Dispatch Controller Development in Hydrogen Multi-Energy Systems
-
$\mu_0$: A Scalable 3D Interaction-Trace World Model
-
Efficient Rationale-based Retrieval: On-policy Distillation from Generative Rerankers based on JEPA
-
Simulating Students' Java Programming Errors with Large Language Models
-
Quantizing Time-Series Models As Dynamical Systems: Trajectory-Based Quantization Sensitivity Score
-
Aligned but Stereotypical? How System Prompts Shape Demographic Bias in LLM-Based Text-to-Image Models
-
From Sorting Algorithms to Scalable Kernels: Bayesian Optimization in High-Dimensional Permutation Spaces
-
Can Post-Training Turn LLMs into Good Medical Coders? An Empirical Study of Generative ICD Coding
-
LoSoNA: A Benchmark for Local Social Norm Adaptation in Group Conversations
-
Toward 360-Degree Indoor Panorama Editing via Tuning-Free Diffusion Model with Refocusing Cross-Attention
-
Neither Parallel Nor Sequential: How DiffusionGemma Actually Commits Tokens
-
Generative Modeling of Bach-Style Symbolic Music: A Comparative Study of Autoregressive, Latent-Variable, and Adversarial Approaches
-
MooMIns -- Monocular 3D Reconstruction and Object Pose Estimation from Multiple Instances
-
From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails
-
The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation
-
Navigating Gigapixel Pathology Images with Large Multimodal Models
-
EyeTheia: A Lightweight and Accessible Eye-Tracking Toolbox
-
HairPort: In-context 3D-aware Hair Import and Transfer for Images
-
C-QUERI: Congressional Questions, Exchanges, and Responses in Institutions Dataset
-
MARD: Mirror-Augmented Reasoning Distillation for Mechanism-Level Drug-Drug Interaction Prediction
-
GEASS: Gated Evidence-Adaptive Selective Caption Trust for Vision-Language Models
-
LLMs as ASP Programmers: Self-Correction Enables Task-Agnostic Nonmonotonic Reasoning
-
The KG-ER Conceptual Schema Language
-
VLADriveBench: Evaluating CoT-Action Relationship in VLA for Autonomous Driving
-
Allure of Craquelure: A Variational-Generative Approach to Crack Detection in Paintings
-
Reliability of Probabilistic Emulation of Physical Systems
-
Entropic Mirror Monte Carlo
-
SciR: A Controllable Benchmark for Scientific Reasoning in LLMs
-
EPIG: Emotion-Based Prompting for Personalised Image Generation
-
ERTS: Adversarial Robustness Testing of Ethical AI via Semantic Perturbation in a Bounded Consequence Space
-
GeoDial: A Multimodal Conversational Tutoring Dataset for Geometry Problem-Solving with Visual Tutor Turns
-
Automated reproducibility assessments in the social and behavioral sciences using large language models
-
Democracy in the Era of Artificial Intelligence
-
Select and Improve: Understanding the Mechanics of Post-Training for Reasoning
-
AgentRivet: an automated system for producing Rivet routines from journal publications
-
HalluJudge: A Reference-Free Hallucination Detection for Context Misalignment in Code Review Automation
-
Meta-Learning Transformers to Improve In-Context Generalization
-
Understanding the Rejection of Fixes Generated by Agentic Pull Requests -- Insights from the AIDev Dataset
-
CodeAlchemy: Synthetic Code Rewriting at Scale
-
Quality Is Not a Safety Proxy Under Quantization
-
Optimal Post-Training Quantization Scales and Where to Find Them
-
Recoverable but Not Stationary:Local Linear Structures in Weights and Activations
-
Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use
-
T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains
-
Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models
-
Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement
-
ASTRA-sim 3.0: Next-Level Distributed Machine Learning Simulations via High-Fidelity GPU and Infrastructure Modeling
-
Do LLMsMakeNeural Distinguishers Wise?
-
When Attribution Patching Lies: Diagnosis and a Second-Order Correction
-
MetaPlate: Counterfactual-Guided RAG-LLM Tool for Personalized Food Recommendation and Hyperglycemia Prevention
-
Local Is Not a Sufficient Privacy Boundary: Governing OS-Integrated On-Device AI
-
Linguistically Augmented Audio Speech Data (LinguAS)
-
Assessing Automated Prompt Injection Attacks in Agentic Environments
-
Machine Learning Methods for Studying Latent Neural Activity Dynamics
-
What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects
-
CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback
-
Moonshine: An Autonomous Mathematical Research Agent Centered on Conjecture Generation
-
Deployment-Time Memorization in Foundation-Model Agents
-
CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs
-
Ethical and Technical Limits of Deepfake Speech Datasets
-
Don't waste SAM
-
A Constrained Natural-Language Interface for Variational Multi-Physics Finite Element Simulations in FEniCS
-
Beyond Point Estimates: Benchmarking Uncertainty Quantification Methods on the AION-1 Astronomical Foundation Model
-
Can You Trust What You See? Human and AI Detection of Synthetic Legal Evidence
-
mllm-shap: A Shapley Value Explainability Platform for Text-Audio Multimodal Large Language Models
-
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
-
Capacity, Not Format: Rethinking Structured Reasoning Failures
-
Bridging Expert Knowledge and Automated Feature Engineering via Self-Evolution
-
The CIFAR Synthetic Evidence Corpus for Detecting AI-Generated Evidence
-
Cross-LLM Consistency in Inference: Evidence from Shared Interactions
-
Vector Space of Cycles
-
Orange Lab: Lowering Barriers to Data Mining through Embedded Interactive Workflows
-
Kunlun: Establishing Scaling Laws for Massive-Scale Recommendation Systems through Unified Architecture Design
-
GIScholarBench: Benchmarking LLM Overconfidence in GIS Research
-
SemDINO: A DINOv3-Driven Network for Cross-Temporal Semantic Alignment in Change Detection
-
Investigating the Histogram Loss in Regression
-
Reconstructing and forecasting disease trajectories of patients with Alzheimer's disease using routine data in resource-constrained settings
-
Reflection in the Dark: Exposing and Escaping the Black Box in Reflective Prompt Optimization
-
From 0-to-1 to 1-to-N: Reproducible Engineering Evidence for MetaAI Recursive Self-Design
-
Where Does the Answer Come From? Benchmarking View-Level Visual Evidence Identification in Multi-View MLLMs for Autonomous Driving
-
C3VD-DEFCOL: A Deformable Colonoscopy Dataset with Time-Resolved 3D Ground Truth and Realistic Appearance
-
Leveraging NeRF-Rendered Images for 3D Gaussian Splatting
-
CP4D: Compositional Physics-aware 4D Scene Generation
-
GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization
-
CATPO: Critique-Augmented Tree Policy Optimization
-
Insertion Based Sequence Generation with Learnable Order Dynamics
-
DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation
-
Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction
-
Learning Quantized Continuous Controllers for Integer Hardware
-
Brain2Text Decoding Model Reveals the Neural Mechanisms of Visual Semantic Processing
-
Implementing Grassroots Logic Programs with Multiagent Transition Systems and AI (Full Version)
-
Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attacks
-
CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models
-
Trajectory Geometry of Transformer Representations Across Layers
-
Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy
-
Steganography Without Modification: Hidden Communication via LLM Seeds
-
LargeMonitor: Monitoring Online Task-Free Continual Learning via Large Pretrained Models
-
Executable World Models for ARC-AGI-3 in the Era of Coding Agents
-
Advancing Mathematics Research with AI-Driven Formal Proof Search
-
When Tabular Foundation Models Meet Strategic Tabular Data: A Prior Alignment Approach
-
You Only Landmark Once: Lightweight U-Net Face Super Resolution with YOLO-World Landmark Heatmaps
-
The Identity Trap in EEG Foundation Models: A Diagnostic Audit
-
ThinkBooster: A Unified Framework for Seamless Test-Time Scaling of LLM Reasoning
-
UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding
-
Exploring Flow-Lenia Universes with a Curiosity-driven AI Scientist: Discovering Diverse Ecosystem Dynamics
-
TSAQA: Time Series Analysis Question And Answering Benchmark
-
SigmaScale: LLM Compression with SVD-based Low-Rank Decomposition and Learned Scaling Matrices
-
Explain Like I'm 5 or Whatever I Choose: Evaluating the Interactive Potential of Language Model Responses
-
When CLIP Sees More, It Fights Back Harder: Multi-View Guided Adaptive Counterattacks for Test-Time Adversarial Robustness
-
Multi-objective optimization and quantum hybridization of equivariant deep learning interatomic potentials
-
Hearing the Unspoken: Language Model Priors for Acoustic Adversarial Attacks
-
It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents
-
A Dynamic Self-Evolving Extraction System
-
Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets
-
Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs
-
TabSwift: An Efficient Tabular Foundation Model with Row-Wise Attention
-
Breaking the Ice: Analyzing Cold Start Latency in vLLM
-
KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026
-
Trio: Learning Time-Series Forecasting with Temporal-Spatial-Sample Attention and Structural Causal Priors
-
WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark
-
M$^3$Exam: Benchmarking Multimodal Memory for Realistic User-Agent Interactions
-
More Capable, Less Cooperative? When LLMs Fail At Zero-Cost Collaboration