Entities · Locations
arXiv
837 articles tagged with this entity.
-
Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety
-
Large Behavior Model: A Promptable Digital Twin of the Retail Customer
-
Towards Understanding Steering Strength
-
CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis
-
An optimal control approach for neural network architecture adaptation with a posteriori error estimation
-
Pelican-VLA 0.5: Attending Before Acting Benefits Generalization
-
Object Search in Partially-Known Environments via LLM-informed Model-based Planning and Prompt Selection
-
Face-trace: Open-Set Attribution and Progressive Discovery of Synthetic Face Generators
-
Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders
-
NonTextual Target Attack
-
RWGBench: Evaluating Scholarly Positioning in Related Work Generation
-
HumAIN: Human-Aware Implicit Social Robot Navigation
-
From Beats to Breaches:How Offensive AI Infers Sensitive User Information from Playlists
-
Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning
-
Predicting LLM Safety Before Release by Simulating Deployment
-
GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model
-
InfraQR: Edge-Placed QR-Inspired Structured Patch Attacks on Infrared Vision-Language Models
-
Thinking Ahead: Foresight Intelligence in MLLMs and World Model
-
CompDiff: Hierarchical Compositional Diffusion for Fair and Zero-Shot Intersectional Medical Image Generation
-
Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection
-
QANTIS: Hardware-Calibrated Sequential POMDP Belief Updates on IBM Heron
-
Optimal Conformal Prediction under Epistemic Uncertainty
-
What Predicts Correctness in Text-to-SQL? A Selective-Prediction Study
-
Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks?
-
Quantifying Frontier LLM Capabilities for Container Sandbox Escape
-
Agent Step Value: Probing the Observer Effect in Black-Box Traces
-
Faithful or Findable? Evaluating LLM-Generated Metadata for RDF Dataset Search
-
Whose fairness? Structural concentration in AI bias research
-
PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation
-
Beyond Static Evaluation: Building Simulation Environments for Scalable Agentic Reinforcement Learning
-
BabyVision: Visual Reasoning Beyond Language
-
ArtisanCAD: An Industrial-Level CAD Agent with Expert-Grounded Knowledge Distillation
-
FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games
-
Black Hole Black Boxes: Numerical Black Hole Metrics via AInstein Neural Networks
-
PerCaM-Health: Personalized Dynamic Causal Graphs for Healthcare Reasoning
-
Reliable Mislabel Detection for Video Capsule Endoscopy Data
-
HumanOmni-Speaker: Identifying Who said What and When
-
Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale
-
Beyond Refusal: A Same-Lineage Study of Aligned and Abliterated LLMs for Vulnerability Analysis
-
An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery
-
Narrative-Centered Emotional Reflection: An Early Prototype for AI-Supported Emotional Self-Reflection
-
Benchmarking KV-Cache Optimizations across Task Quality and System Performance for Long-Context Serving
-
RPAM: A Principled Metric for Evaluating Associations in Language Models with High Predictive Validity in Downstream Outputs
-
When Should LLMs Search? Counterfactual Supervision for Search Routing
-
FourTune: Towards Fully 4-Bit Efficient Post-Training for Diffusion Models
-
KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking
-
Towards Interpretable Foundation Models for Retinal Fundus Images
-
AlayaWorld: Long-Horizon and Playable Video World Generation
-
Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade
-
aiAuthZ: Off-Host, Identity-Bound Authorization for AI Agents
-
Mitigating Errors in LLM-Generated Web API Invocations via Retrieval-Augmented Generation and Constrained Decoding
-
Reduced NEXI protocol for the quantification of human gray matter microstructure on the Connectome 2.0 scanner
-
Detoxify: A framework for abusive text transformation using LLMs
-
Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design
-
Spider 2.0-AIFunc: Extending Real-World Text-to-SQL to AI-Native SQL Workflows
-
WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS
-
Agentic Artificial Intelligence for Multistage Physics Experiments at a Large-Scale User Facility Particle Accelerator
-
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
-
EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization
-
RustMizan: A Compilable, Contamination-Aware Benchmarking Framework for Rust Vulnerabilities
-
Equivalence of Context and Parameter Updates in Modern Transformer Blocks
-
Effective Distillation to Hybrid xLSTM Architectures
-
Reading Between the Dots: Decoding Hidden Computation across Filler Tokens
-
Decomposed Prompting Does Not Fix Knowledge Gaps, But Helps Models Say "I Don't Know"
-
Findings of the Fifth Shared Task on Multilingual Coreference Resolution: Expanding Datasets for Long-Range Entities
-
Differentiate the Evaluator, Not the Program: An Efficient Runtime Representation for Neuro-Symbolic Learning
-
Knowledge-Centric Information Systems
-
When Does Small Data Work? Accuracy and Efficiency Trade-offs Between Tabular Foundation Models and Conventional Methods for Crowd-State Classification at Hajj and Umrah
-
Open Problem: Is Interaction Necessary for Order-Optimal 1-bit Mean Estimation?
-
SMOCS: A Streaming Framework for Simplified Deployment, Monitoring, and Optimization of ML Systems in Production
-
LeukocyteCount: Automatic Identification and Counting for leukocytes using Deep Learning
-
R3D: Quantitative 3D Spatial Reasoning for Egocentric Wearables
-
Latent Programming Horizons in Coding Agents
-
Learning Flexible Generalization in Video Quality Assessment by Bringing Device and Viewing Condition Distributions
-
GALOSH: Blind, Training-Free Denoising of Raw Bayer and sRGB Images by Parallel-Friendly Local Shrinkage
-
SIMPLER: Efficient Foundation Model Adaptation via Similarity-Guided Layer Pruning for Earth Observation
-
When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents
-
Programming over Thinking: Efficient and Robust Multi-Constraint Planning
-
In-span learning: adapting reduced-order models using their own predictions
-
SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
-
A Digital Twin Framework for Traffic-Aware UAV Pavement Monitoring in Open-Traffic Conditions
-
Latent Visual Cache for Video Reasoning
-
G$^2$TAM: Geometry Grounded Track Anything Model
-
PreSIST: Vision-Language-Informed Object Persistence Prediction in Open-World Scenes
-
RAF: Reliability-Aware Fusion of Camera, LiDAR, and 4D RADAR for Robust 3D Object Detection in Adverse Weather
-
ProbeLogits: Kernel-Level LLM Inference Primitives for AI-Native Operating Systems
-
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry
-
Evaluating and Understanding Model Editing for Medical Vision Language Models
-
Metronome: Bound the Cache, Keep the Beat for Real-Time Interaction Model Serving
-
Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models
-
Where do LLMs Fall Short in CBT-Guided Affective Reasoning?
-
OpenGlass: A Sensing-Computing Split Architecture for Local MLLM-Driven Real-Time Visual Assistance
-
Builder, Defender, Breaker: The Case Against Removing the Human from the AI-Driven Security Lifecycle
-
Semantic Segmentation-Driven Image-Level Diagnosis of Liver Cancers in Hematoxylin and Eosin Histopathology Images
-
Scalable Maximal Frequent Episode Mining with Desbordante
-
Seeing Once is Enough? Online Geometry-Aware Token Pruning for 3D Question Answering
-
Tightening the Score Matching Gap for Diffusion Models
-
TrendFact: A Benchmark Towards Hotspot Perception in Automatic Fact-Checking
-
Attributing Emergence in Million-Agent Systems
-
Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI
-
Which Algorithm Specification Formats Help Language Models Implement Machine Learning Algorithms?
-
Target-Aware Interaction-Guided Reinforcement Learning for Black-Box Node Injection Attacks on Graph Neural Networks
-
NKI-Agent: Domain-Specific Fine-Tuning and Agentic Tool Use for Neuron Kernel Generation
-
The Map Behind the Flow: Finite-Step Gradient Descent as a Dynamical System
-
Mechanism-level routing failure in LLMs over Lean-verified algebraic structures
-
RABBiT: Rapidly adaptive BOLD foundation model via brain-tuning for accurate zero-shot and few-shot prediction of speech-elicited responses in the brain
-
ChatImage: Navigating Long-Form LLM Answers through Interactive Images
-
TimeThink: Reasoning with Time for Video LLMs
-
Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs
-
TestMate: Test-Time Domain Adaptation Aided by Lightweight Vision Foundation Model
-
Last-Meter Precision Navigation for UAVs: A Diffusion-Refined Aerial Visual Servoing Approach
-
EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions
-
An Exploration of Agentic Information Fusion for Test Maintenance Prediction
-
Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents
-
How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Kr\"ugel, and Uhl (2025)
-
When Simpler Is Better: Evaluating Translation Pipelines for Medieval Latin Manuscripts
-
Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure
-
Seduced by the Narrative: Assessing Rule Adherence in Semi-Open Textual Sandboxes
-
Improving LLMs via Validator-to-Generator Alignment
-
Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs
-
UNITY: Attention Flow Networks for Adaptive Conditioning in Diffusion
-
E-TraMamba: A New Paradigm for Efficient Long-Term 3D Feature Tracking with Event Cameras
-
AI Wizards at EXIST 2026: Hierarchical Soft-Label Learning for Multimodal Sexism Identification in Memes
-
Large-scale dataset of automatically classified rhetorical sections in scientific papers
-
Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?
-
MedCalc-Pro: Solving Complex Medical Calculations with LLM Agents
-
Human-Centric Reflective Architecture for Human-AI Collaborative Decision-Making
-
Detecting Answer-Driven Reasoning in LLM-Based Educational Tutors via Truncated Chain-of-Thought Auditing
-
EMPURPLE: A Free Lunch for Diffusion Distillation based on the Information Bottleneck
-
Beyond Post-Quantization: Native Hash Learning with a Dedicated HASH Token
-
Your Agent's Memories Are Not Its Own: Forged Reasoning Attacks on LLM Agent Memory and Defenses
-
Walma: Learning to See Memory Corruption in WebAssembly
-
OpenTinker: Separating Concerns in Agentic Reinforcement Learning
-
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
-
LangLoc: "Tell Me What You See"
-
SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion
-
Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
-
VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
-
DualView: Preventing Indirect Prompt Injection in Personal AI Agents
-
MentalThink: Shaping Thoughts in Mental SVG World
-
Prior Bias in Vision Language Models on UML Diagram Interpretation
-
Enhanced Feature Extraction for IoT Network Intrusion Detection Using GNNs and KAN
-
Unified Audio Intelligence Without Regressing on Text Intelligence
-
GenHOI: Generalized Hand-Object Pose Estimation with Occlusion Awareness
-
TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution
-
HAL: Inducing Human-likeness in LLMs with Alignment
-
Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment
-
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
-
Hyperloop Transformers
-
Probabilistic Low-Voltage Peak Load Forecasting with Time Series Foundation Models Evaluated on Application-Oriented Metrics
-
Epistemic Goggles: A Pretrained Module that Induces an Epistemic Frame via Gradient Editing
-
Meta-Benchmarks for Financial-Services LLM Evaluation
-
ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair
-
Collaborative Disagreement Resolution for Scalable Oversight
-
WorldSample: Closed-loop Real-robot RL with World Modelling
-
A global optimization SAR image segmentation model can be easily transformed to a general ROF denoising model
-
kNNGuard: Turning LLM Hidden Activations into a Training-Free Configurable Guardrail
-
LLMs as Teaching Assistants for Mathematics Exam Grading: Reliability, and Practical Usability
-
Single-Channel EEG-Based Cognitive Load Assessment in Online Learning: A Hybrid Deep Learning Approach
-
Has This Checkpoint Been Abliterated? A Two-Signal Audit and Its Failure Map
-
BRIDGE: Predicting Human Task Completion Time From Model Performance
-
Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
-
Beyond Skepticism: Evaluating LLMs Pedagogical Intent Reasoning with the Adaptive Pedagogical Vigilance Framework
-
PARTREP: Learning What to Repeat for Decoder-only LLMs
-
AlienLM: Alienization of Language for API-Boundary Privacy in Black-Box LLMs
-
Multi-THuMBS: Multi-person Tracking of 3D Human Meshes Beyond Video Shots
-
GeoMix: Descriptor-Free Visual Localization via Global Context and Multi-Detector Training
-
Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows
-
World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments
-
COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows
-
Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification
-
MMIR-TCM: Memory-Integrated Multimodal Inference and Retrieval for TCM Clinical Decision Support
-
ContextNest: Verifiable Context Governance for Autonomous AI Agent
-
Safeguarding LLM Agents from Misalignment through Provenance Analysis
-
Safety Targeted Embedding Exploit via Refinement
-
Rethinking Complexity Metrics for LLM-Integrated Applications: Beyond Source Code
-
Rank-Then-Act: Reward-Free Control from Frame-Order Progress
-
Gaming Consensus: Coordinated Manipulation in Crowdsourced Fact-Checking
-
ESC: Emotional Self-Correction for Reliable Vision-Language Models
-
What Types of Human-AI Teams Exist?
-
HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety
-
GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation
-
TiRex-2: Generalizing TiRex to Multivariate Data and Streaming
-
EPC: A Standardized Protocol for Measuring Evaluator Preference Dynamics in LLM Agent Systems
-
Learning dynamical systems from noisy data with Weak-form Kernel Ridge Regression
-
GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models
-
GameDevBench: Evaluating Agentic Capabilities Through Game Development
-
Utilizing Earth Foundation Models to Enhance the Simulation Performance of Hydrological Models with AlphaEarth Embeddings
-
Planning over MAPF Agent Dependencies via Multi-Dependency PIBT
-
StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning
-
From Signals to Structure: How Memory Architecture Drives Language Emergence in LLM Agents
-
Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising
-
GKDT: General Keypoint Detection Transformer
-
Towards High-Resolution Visual Perception via Hierarchical Entity Exploration
-
From "Strings" to "Things" for Personal Knowledge Graphs: Evaluating LLM Triple Extraction for Recommendation Systems
-
NeuroCogMap Reveals Cognitive Organization of Large Language Models
-
Autonomous Scientific Discovery via Iterative Meta-Reflection
-
Measuring the Gap Between Human and LLM Research Ideas
-
Crystalite: A Lightweight Transformer for Efficient Crystal Modeling
-
Explaining Tabular Foundation Model Differences Through Meta-Features
-
Comparing Large Language Models on Scrum Certification-Style Questions: Accuracy, Stability, and Error Patterns
-
The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology
-
A Text-Steerable Instrument for Sketching Procedural Soundscapes via Language Models
-
A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models
-
PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents
-
Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences
-
Structural Enforcement of Statistical Rigor in AI-Driven Discovery: A Functional Architecture
-
Generating consensus and dissent on massive discussion platforms with a semantic-vector model
-
Position: Collaborative Agentic AI Needs Interoperability Across Ecosystems
-
LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent
-
Paper2Rebuttal: A Multi-Agent Framework for Transparent Author Response Assistance
-
Disentangling Reasoning Logic to Resolve Explicit Knowledge Conflicts
-
FLARE-AI: Flaw Reporting for AI
-
Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support
-
Toward AI-Resilient Assessment in Computer Science Courses in an AI-Native World
-
Large Databases Need Small, Open-Weight Language Models
-
A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols
-
PolicyGuard: From Organizational Policies to Neuro-SymbolicCompliance Review Engines
-
Bridging Local Observation and Global Simulation in Closed-Loop Traffic Modeling
-
Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles
-
SMART: When is it Actually Worth Expanding a Speculative Tree?
-
Multilingual Polarization Detection Using Transformer-Based Models with Class Weighting and Threshold Tuning
-
Measuring Judgment Quality in Natural-Language Explanations: Evidence from Forecasting Tournaments
-
REMSA: Foundation Model Selection for Remote Sensing via a Constraint-Aware Agent
-
Efficient Public Verification of Private ML via Regularization
-
When the Database Fails: Prompting LLM Dialogue Agents for Safe Recovery in Task-Oriented Dialogue
-
PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks
-
Temporal Preservation over Processing: Diagnosing and Designing Spatiotemporal Single-Stage Video Detectors
-
FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents
-
ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration
-
Agentic Tool Use in Large Language Models
-
Mitigating Batch Effects in Histopathology via Language-Mediated Robust Embedding Generation
-
MIRROR: Aligning Semantic Relations from Language to Image via Gromov--Wasserstein
-
Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization
-
Predicting Effects, Missing Distributions: Evaluating LLMs as Human Behavior Simulators in Operations Management
-
Overcoming Dependent Censoring in the Evaluation of Survival Models
-
Primary ICD Category Prediction using LLM-based Probing
-
ManimAgent: Self-Evolving Multimodal Agents for Visual Education
-
LLMography: Transforming Human-AI Conversations into Traceability, Oversight, and Auditability Indicators
-
MotionAtlas: Detailed Region Captioning for Motion-Centric Videos
-
ArchesClimate: Probabilistic Decadal Ensemble Generation With Flow Matching
-
Rigel: Self-Distilled Score Adaptation for Image and Video Captioning Evaluation
-
wav2VOT: Automatic estimation of voice onset time, closure duration, and burst realisation with wav2vec2
-
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks
-
Labeling Training Data for Entity Matching Using Large Language Models
-
Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Models
-
Conversational Domain Adaptation of IndicTrans2 across 21 Indic Languages via Experience Replay and Model Soups
-
SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models
-
Timesteps of Mamba Align with Human Reading Times
-
LLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive Dashboard
-
Grounding LLM Reasoning under Incomplete Graph Evidence
-
FAIL: Flow Matching Adversarial Imitation Learning for Image Generation
-
Value-Action Alignment in Large Language Models under Privacy-Prosocial Conflict
-
On the Nonlinearity of Learning Rate Scaling for LLM Training
-
Beyond IID: How General Are Tabular Foundation Models, Really?
-
TextClusterLab: An Integrated Framework for Reliable Text Clustering Studies
-
Structural Certification for Reliable Physical Design with Language Models
-
Entity Binding Failures in Tool-Augmented Agents
-
Trust Your Instincts: Confidence-Driven Test-Time RL for Vision-Language-Action Models
-
CLMASP: Coupling Large Language Models with Answer Set Programming for Robotic Task Planning
-
NEURON: A Neuro-symbolic System for Grounded Clinical Explainability
-
CAPTCHA Solving for Native GUI Agents: Automated Reasoning-Action Data Generation and Self-Corrective Training
-
Obliviate: Erasing Concepts from Autoregressive Image Generation Models
-
A Classifier-Agnostic Zero-Shot Adversarial Attack Detection via CLIP
-
Learning to Segment Liquids in Real-world Images
-
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
-
Beyond the Reranker: Do RAG Retrieval Enhancements Help Once a Strong Reranker Is Present?
-
Conversational Query Engine for Mixed-Modality Heterogeneous Enterprise Data Sources
-
Building to the Test: Coding Agents Deliver What You Check, Not What You Requested
-
UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception
-
Capability Gates Are Not Authorization: Confused-Deputy Failures in LLM Agent Frameworks
-
Proteus: Automated Adversarial Robustness Testing for Audio Deepfake Detectors
-
StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley
-
Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks
-
EcoVideo: Entropy-Orchestrated Video Generation Paradigm in Cloud-Edge Dynamics
-
Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model
-
The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction
-
Extrapolating from Regularised Solutions for Solving Ill-Conditioned Linear Systems in Machine Learning
-
When Web Agents Finish but Still Fail: Reproducible Triggers and Trace Diagnostics for Parallel Web Exploration
-
IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations
-
Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing
-
AI Trading's Alpha Singularity: Emergent Market Reasoning through Agent-to-Agent Self-Evolution
-
The CRISTAL Method: Neurosymbolic analysis from AI-synthesized world models
-
Rethinking Generative Reconstruction Attacks against Graph Neural Network Models
-
Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents
-
Vision-Language-Action Models: Experimental Insights from a Real-World UR5 Platform
-
The Digital Afterlife of Empires: Four Language Models Converge on the Same Imperial Cartography of Writing
-
Detecting Clinical Hallucinations in LVLMs via Counterfactual Visual Grounding Uncertainty
-
Em-ergence of the em-dash: a population-level rise in em-dash frequency in medRxiv preprints at the dawn of the large-language-model era
-
SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics
-
ML-Powered LDAP Reconnaissance Detection using Weak Supervision
-
Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?
-
Hierarchical Experimentalist Agents
-
Dynamic Parsing and Updating Natural Language Specification using VLMs for Robust Vision-Language Tracking
-
fev-bench: A Realistic Benchmark for Time Series Forecasting
-
Latent Bridges for Multi-Table Question Answering
-
How Far Can You Get Without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation
-
TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents
-
Categorize Early, Integrate Late: Divergent Processing Strategies in Automatic Speech Recognition
-
Enhanced Diffusion Sampling: Efficient Rare Event Sampling and Free Energy Calculation with Diffusion Models
-
DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information
-
SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions
-
Diversity is the Strength of the AI Crowd
-
SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing
-
Foundation vs. Specialized Models: Evaluating Catastrophic Forgetting in Continual Time Series Forecasting
-
Accelerating Hierarchical Sparse Predictive Coding with Hybrid Amortized Inference
-
Neural Image Space Tessellation effect
-
StableMotion: One-Step Motion Estimation with Diffusion Prior
-
Unified Zero-Shot Time Series Forecasting: A Darts Foundation
-
Auto-Configuring Scientific Simulators with Lightweight Coding-Agent Adapters
-
DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain
-
GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models
-
Multimodal Evaluator Preference Collapse: Cross-Modal Coupling in Self-Evolving Agents
-
Textual Belief States for World Models: Identifiable Representation Learning Under Strict Mediation
-
OperatorSHAP: Fast and Accurate Shapley Value Estimation for Neural Operators
-
Self-Stigma Is Not a Monolith, but Generic Empathy Is: Persona-Conditioned LLM Support for People Who Use Drugs
-
ReWorld: Learning Better Representations for World Action Models
-
Radar Guided Camera Verification for Automatic Emergency Braking Rethinking Object Detection in Radar Camera Fusion
-
Are Time-Series Foundation Models Ready for E-Nose Data? An Empirical Assessment of Their Embeddings
-
ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents
-
Hybrid Fact-Checking that Integrates Knowledge Graphs, Large Language Models, and Search-Based Retrieval Agents Improves Interpretable Claim Verification
-
LieSolver: PDE-Constrained Learning for IBVPs via Lie Symmetries
-
Scene and Human in One World: Reconstruction in a Feedforward Pass
-
Optimizing Teacher-Student Partitioning for Scalable Knowledge Distillation on HPC Systems
-
RSD: Moving Local Triangular Charts for Auditing Language-Model Hidden States
-
Contagion Networks: Evaluator Preference Propagation in Multi-Agent LLM Systems
-
When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search
-
ToE: A Hierarchical and Explainable Claim Verification Framework with Dynamic Multi-source Evidence Retrieval and Aggregation
-
CalBrief: A Pilot Diagnostic Benchmark for Evidence-Calibrated Scientific Briefing with Large Language Models
-
PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration
-
Deployment-Side Adaptiveness in Multi-Horizon Volatility Forecasting
-
Does Aurora Encode Atmospheric Structure? Latent Regime Analysis and Attribution
-
RecallRisk-BERT: A Multi-Task Framework for Post-Report Medical Device Recall Triage
-
A3C3: AI Algorithm and Accelerator Co-design, Co-search, and Co-generation
-
Paved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training Regimes
-
How Good Can Linear Models Be for Time-Series Forecasting?
-
A Pipeline for Generating Longitudinal Synthetic Clinical Notes Using Large Language Models
-
Nemotron-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context
-
Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness
-
Divergent Recommendations, Convergent Diagnoses: Cross-Provider Failure-Mode Convergence in AI Commercial Recommendation
-
Beyond Logical Forms: LLM-Extracted Patterns for Fallacy Classification
-
Empirical Software Engineering TerraProbe: A Layered-Oracle Framework for Detecting Deceptive Fixes in LLM-Assisted Terraform
-
Hybrid privacy-aware semantic search: SVD-truncated document geometry and CKKS-encrypted query reranking under a restricted threat model
-
Learning State-Tracking from Code Using Linear RNNs
-
PrivacyBench: Privacy Isn't Free in Hybrid Privacy-Preserving Vision Systems
-
A Generalization Theory for JEPA-Based World Models
-
Closing the Loop to Discover Psychological Theories with an Automated Cognitive Scientist
-
An Empirical Study of LLM-Generated Specifications for VeriFast
-
The Open Source Economic Index of AI Adoption and Capability
-
S2P-Net: A Spectral-Spatial Polar Network for Rotation-Invariant Object Recognition in Low-Data Regimes
-
HierBias: Context-Conditioned Hierarchical Media Bias Detection with Multi-Task Type Classification
-
SocialPersona: Benchmarking Personalized Profiling and Response with Multimodal Social-Media Context
-
Do Image Editing Models Understand Lighting?
-
PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models
-
Boundary-Aware Context Grounding for A Low-Channel EEG Agent
-
Beyond Surface Forms: A Comprehensive, Mechanism-Oriented Taxonomy of Indirect Linguistic Encoding for LLM-Based Coded Language Detection
-
EGG: An Expert-Guided Agent Framework for Kernel Generation
-
Vulnerability of Natural Language Classifiers to Evolutionary Generated Adversarial Text
-
Thinking Like a Scientist? A Structural Study of LLM-Generated Research Methods
-
Extracting Neural Materials from Multi-view Images
-
CyberChainBench: Can AI Agents Secure Smart Contracts Against Real-World On-Chain Vulnerabilities?
-
AXLE: A Cloud Infrastructure for Lean 4 Theorem Proving Utilities
-
6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models
-
Accelerating Returns and the Qualitative Engine for Science
-
Autoregressive Boltzmann Generators
-
SciFig: Towards Automating Editable Figure Generation for Scientific Papers
-
Limited Reference, Reliable Generation: A Two-Component Framework for Tabular Data Generation in Low-Data Regimes
-
Use What You Know: Causal Foundation Models with Partial Graphs
-
Kolmogorov Arnold networks (KAN) for aerodynamic prediction: a comparison with MLPs and GNNs
-
Ramanujan Graph Rewiring with Non Negative Resistance Curvature
-
Membox: Weaving Topic Continuity into Long-Range Memory for LLM Agents
-
Benchmarking Vision-Language Models for Microscopic Plant Image Understanding
-
ACT-JEPA: Novel Joint-Embedding Predictive Architecture for Efficient Policy Representation Learning
-
A Probabilistic Framework for LLM-Based Model Discovery
-
Fox in the Henhouse: Supply-Chain Backdoor Attacks Against Reinforcement Learning
-
Kuramoto Oscillatory Phase Encoding: Neuro-inspired Synchronization for Improved Learning Efficiency
-
TensorLDM: A Component-Wise Latent Diffusion Model for Volumetric DTI Reconstruction from Sparse DWIs
-
Steering Vision-Language Models with Joint Sparse Autoencoders
-
A Controlled Study of CLIP-Based Body-Scene Fusion for Emotion Recognition in Context
-
C3-Bench: A Context-Aware Change Captioning Benchmark
-
Flexible Gravitational-Wave Parameter Estimation with Transformers
-
HG-Bench: A Benchmark for Multi-Page Handwritten Answer-Region Grounding in Automated Homework Assessment
-
Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs
-
OracleAnalyser: Analysing Implicit Semantics of Oracle Bone Scripts through MLLMs with Post-training
-
RoboAtlas: Contextual Active SLAM
-
Benchmarking Deep Learning Models for Laryngeal Cancer Staging Using the LaryngealCT Dataset
-
Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding
-
LLM-Based Scientific Peer Review: Methods, Benchmarks, and Reliability Challenges
-
Story Operators: Decomposing the Original $\to$ Sequel Transformation in Embedding Space
-
Small edits, large models: How Wikipedia advocacy shapes LLM values
-
Graph-Based Phonetic Error Correction of Noisy ASR
-
The Warrant Gap: Claim-Conditioned Re-scoring for Fact-Checking
-
PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D Garments
-
AI Fiction in the Wild
-
Latent Visual States for Efficient Multimodal Reasoning
-
DramaDirector: Geometry-Guided Short Drama Generation
-
Average Rankings Mask Per-Subject Optimality: A Friedman-Nemenyi Benchmark of EEG Motor-Imagery BCI Decoders
-
Paying to Know: Micro-Transaction Markets for Verified Product Information in Agentic E-Commerce
-
BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents
-
Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment
-
Which Spaces can be Embedded in $L_p$-type Reproducing Kernel Banach Space? A Characterization via Metric Entropy
-
Impatient Bandits: Optimizing for the Long-Term Without Delay
-
Dimensionality Reduction of QAOA Parameter Space with Kernel PCA for Max-Cut
-
HyMaTE: A Hybrid Mamba and Transformer Model for EHR Representation Learning
-
Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable
-
Prob-BBDM: a Probabilistic Brownian Bridge Diffusion Model for MRI sequence image-to-image translation
-
VeriPilot: An LLM-Powered Verilog Debugging Framework
-
When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs
-
AdversaBench: Automated LLM Red-Teaming with Multi-Judge Confirmation and Cross-Model Transferability
-
Lightweight Transformer Models for On-Device Fault Detection: A Benchmark Study on Resource-Constrained Deployment
-
SP-Mind: An Autonomous Reasoning Agent for Spatial Proteomics Analysis
-
LemonHarness Technical Report
-
Aligning Audio Captions with Human Preferences
-
OpenThoughts-Agent: Data Recipes for Agentic Models
-
The Measurable Majority
-
When AI Meets Finance (StockAgent): Large Language Model-based Stock Trading in Simulated Real-world Environments
-
MassSpecGym in the Wild: Uncovering and Correcting Evaluation Pitfalls in AI-Driven Molecule Discovery
-
Holo-World: Unified Camera, Object and Weather Control for Video World Model
-
A High-Resolution Landscape Dataset for Concept-Based XAI With Application to Species Distribution Models
-
Does Head Pose Correction Improve Biometric Facial Recognition?
-
GEMS: Geometric Constraints Enable Multi-Semantic Superposition in LLMs
-
Self-Preference Is Weak or Absent in Verifiable Instruction-Following Revision: A Four-Model Test Under Genuine Authorship
-
Capturing Intransitive Dominance in Tennis Forecasting: A Graph Neural Network Approach
-
Pseudo-Feature Padding: A Lightweight Defense Against False Data Injection in Power Grids
-
Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems
-
Source-Grounded Data Generation for Text-to-JSON Learning
-
StylisticBias: A Few Human Visual Cues Drive Most Social Biases in MLLMs
-
SAFE-Cascade: Cost-Adaptive Vision-Language Routing for Chart Question Answering
-
Through the PRISM: Preference Representation in Intermediate States of Video Diffusion Models
-
ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis
-
Med-R2: Perception and Reflection-driven Complex Reasoning for Medical Report Generation
-
Exposing the Unsaid: Visualizing Hidden LLM Bias through Stochastic Path Aggregation
-
eCNNTO: A Highly Generalizable ConvNet for Accelerating Topology Optimization
-
BIM-Edit: Benchmarking Large Language Models for IFC-Based Building Information Modeling
-
Physical Atari: A Robust and Accessible Platform for Real-time Reinforcement Learning on Robots
-
Cost-Optimal LLM Routing with Limited User Feedback under User Satisfaction Guarantees
-
Data Standards for Humanoid Robotics: The Missing Infrastructure for Physical AI
-
Measuring Biological Capabilities and Risks of AI Agents
-
Bidirectional Tutoring for Developmental Motor Learning in Robots: Co-Developed Interaction Dynamics Support Stable Learning
-
CREDENCE: Claim Reduction for Decomposition & Enhanced Credibility -- Semantic Metrics and Convergence Analysis
-
Dual-Agent Framework for Cross-Model Verified Translation of Natural-Language Protocols into Robotic Laboratory Platform
-
Creativity Reconsidered: Generative AI and the Problem of Intentional Agency
-
Multi-View Decompilation for LLM-Based Malware Classification
-
Wisdom of Committee: Diverse Distillation from Large Foundation Models and Domain Experts
-
Simulation of Language Evolution under Regulated Social Media Platforms: A Synergistic Approach of Large Language Models and Genetic Algorithms
-
From Construction to Injection: Edit-Based Fingerprints for Large Language Models
-
FM-Agent: Scaling Formal Methods to Large Systems via LLM-Based Hoare-Style Reasoning
-
CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation
-
Self-Adaptive Scale Handling for Forecasting Time Series with Scale Heterogeneity
-
Your Mouse and Eyes Secretly Leak Your Preference: LLM Alignment using Implicit Feedback from Users
-
DeXposure-Claw: An Agentic System for DeFi Risk Supervision
-
Process-Verified Reinforcement Learning for Theorem Proving via Lean
-
Context-Aware Hierarchical Bayesian Modeling of IVF Laboratory Environmental Conditions
-
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models
-
Training-Free Metrics for Synthetic Object Detection Data: A Proxy for Detector Performance
-
Benchmarking Agentic Review Systems
-
GLARE: A Natural Language Interface for Querying Global Explanations
-
MetaResearcher: Scaling Deep Research via Self-Reflective Reinforcement Learning in Adversarial Virtual Environments
-
Multi-Head Attention-Based Feature Extractor Integration with Soft Actor-Critic for Porosity Prediction and Process Parameter Optimization in Additive Manufacturing
-
Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts
-
Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation
-
Towards Scalable Customization and Deployment of Multi-Agent Systems for Enterprise Applications
-
DINO-Med3D: Bridging Dimension and Domain Gaps in Volumetric Segmentation via Progressive Adaptation
-
JourneyFormer: Encoding Airbnb Guest Journey with Sequence Modeling
-
Explaining Attention with Program Synthesis
-
PRISM: A 3D Probabilistic Neural Representation for Interpretable Shape Modeling
-
SwitchBraidNet: Quantisation-Aware Lightweight Architecture for Hybrid Brain-Computer Interface
-
As Easy as Rocket Science: Assessing the Ability of Large Language Models to Interpret Negation in Figurative Language
-
Generating Natural and Expressive Robot Gestures through Iterative Reinforcement Learning with Human Feedback using LLMs
-
MCompassRAG: Topic Metadata as a Semantic Compass for Paragraph-Level Retrieval
-
Machine Unlearning for the XGBoost Model with Network Intrusion Datasets
-
Transformer Geometry Observatory TGO-I: Spectral Geometry Observatory
-
PEC-Home: Interpretation of Progressively Elliptical Commands in Smart Homes
-
From Specification to Execution: AI Assisted Scientific Workflow Management
-
RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models
-
Code-Augur: Agentic Vulnerability Detection via Specification Inference
-
AI-Driven Assessment of Human Tutors: Linking Training Performance to Real-Life Practice
-
Dynamic In-Group Persona Generation for Enhancing Human-AI Rapport
-
G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom Alignment
-
Beyond Tokenization: Direct Timestep Embedding and Contrastive Alignment for Time-Series Question Answering
-
Enhancing CVRP Solver through LLM-driven Automatic Heuristic Design
-
SAMA: Semantic Anchor-aligned Augmentation for Unified Low-Resource Multimodal Information Extraction
-
UniTemp: Unlocking Video Generation in Any Temporal Order via Bidirectional Distillation
-
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
-
Physics-IQ Verified
-
Measurement noise limits the advantage of nonlinear models over linear models in biomedical prediction
-
OpenAnt: LLM-Powered Vulnerability Discovery Through Code Decomposition, Adversarial Verification, and Dynamic Testing
-
Riemannian MeanFlow for One-Step Generation on Manifolds
-
Fluently Lying: Adversarial Robustness Can Be Substrate-Dependent
-
EgoCS-400K: An Egocentric Gameplay Dataset for World Models
-
Quantifying Consistency in LLM Logical Reasoning via Structural Uncertainty
-
FinAcumen: Financial Multimodal Reasoning via Self-Evolving Experience Memory Harness
-
Using Cognitive Models to Improve Language Model Simulation of Human Persuasion Games
-
Surveying GenAI-based Automation in Printed Circuit Board Design and Test
-
Visored: A Controlled-Natural-Language Prover for LLM-Generated Mathematics
-
CIAN: Multi-Stage Framework for Event-Enriched Image Captioning via Retrieval-Augmented Generation
-
STAR: SpatioTemporal Adaptive Reward Allocation for Text-to-Image RL Post-Training
-
From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning
-
A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models
-
Conformal Path Reasoning: Trustworthy Knowledge Graph Question Answering via Path-Level Calibration
-
Environment-Grounded Automated Prompt Optimization for LLM Game Agents
-
Evaluating Open-Source LLMs for Multi-Label ATT&CK Technique Classification on CTI Reports
-
Clarify Before You Draw: Proactive Agents for Robust Text-to-CAD Generation
-
Evaluating Intersectional Fairness across Clinical Machine Learning Use Cases using Fairlogue and the All of Us Research Program
-
Multi-Source Cybersecurity Logs: An ATT&CK-Labeled Dataset and SLM Evaluation
-
Bridging Modality Disconnect in Self-Reflection via Closed-Loop Visually Grounded Verification
-
When the Next Step Is Not One Step: Distribution-Aware Execution Modeling for Concurrent Go Programs
-
TuneAhead: Predicting Fine-tuning Performance Before Full Training Begins
-
Handling Feature Heterogeneity with Learnable Graph Patches
-
Flux-Guard: Facial Identity Protection using diffusion models
-
Know Thy Reasoner: Not All Language Models Explore Alike
-
Planning with the Views
-
SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
-
GeoDisaster: Benchmarking Orchestrated Agents for Operational Disaster Geo-Intelligence
-
Complex Layout Classification in the Wild: A Low-Resource Approach with Layout-Preserving Augmentations
-
Million-scale multimodal pollen microscopy with expert-guided foundation models
-
Seeing Is Not Screening: Multimodal Hidden Instruction Attacks on Agent Skill Scanners
-
ProCUA-SFT Technical Report
-
GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
-
MLLP-VRAIN UPV system for the IWSLT 2026 Simultaneous Speech Translation task
-
Bridging Functional Correctness and Runtime Efficiency Gaps in LLM-Based Code Translation
-
Self-Generated Error Training for Token Editing in Diffusion Language Models
-
RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills
-
Vision-language models for chest radiography do not always need the image
-
EngTrace: A Symbolic Benchmark for Verifiable Process Supervision of Engineering Reasoning
-
Surrogate Assisted Pedestrian Protection Design via a Foundation Model Orchestrated Workflow
-
OpenLID-v3: Improving the Precision of Closely Related Language Identification -- An Experience Report
-
Atlas: Orchestrating Heterogeneous Models and Tools for Multi-Domain Complex Reasoning
-
Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices
-
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
-
Treatment Response Optimized Clinical Decision Support AI System via Digital Twin Simulation
-
FllumaOne: A Code-Native Multimodal CAD Dataset with Executable Programs and Kernel-Validated Feature Histories
-
IsabeLLM: Automated Theorem Proving Applied to Formally Verifying Consensus
-
PreAct: Computer-Using Agents that Get Faster on Repeated Tasks
-
Agentic AI-based Framework for Mitigating Premature Diagnostic Handoff and Silent Hallucination in Healthcare Applications
-
Prefill/Decode-Aware Evaluation of LLM Inference on Emerging AI Accelerators
-
Attribute Inference from Interactive Targeted Ads
-
Agent trajectories as programs: fingerprinting and programming coding-agent behavior
-
ActiveSAM: Image-Conditional Class Pruning for Fast and Accurate Open-Vocabulary Segmentation
-
A Biased Nonnegative Block Term Tensor Decomposition Model for Dynamic QoS Prediction
-
Generative Molecular Design with Steerable and Granular Synthesizability Control
-
MacrOData: New Benchmarks of Thousands of Datasets for Tabular Outlier Detection
-
Unifying Post-hoc Explanations of Knowledge Graph Completions
-
The Machine Learning Approach to Moment Closure Relations for Plasma: A Review
-
Is My Vision-Language Data in Your AI? Membership Inference Test (MINT) Demo 2
-
SPARK: Spatial Policy-driven Adaptive Reinforcement learning for Knowledge distillation
-
Comparing Human Gaze and Vision-Language Model Attention in Safety-Relevant Environments
-
Conditional Multi-Event Temporal Grounding in Long-Form Video
-
EcoBin: A Two-Stage Deep Convolutional Neural Network for Contamination-Aware Waste Classification
-
Domain-Guided Prompting of the Segment Anything Model for Seismic Interpretation: The Role of Attributes, Visualization, and Hybrid Prompts
-
VinQA: Visual Elements Interleaved Long-form Answer Generation for Real-World Multimodal Document QA
-
VisualClaw: A Real-Time, Personalized Agent for the Physical World
-
Uncertainty Quality of VGGT: An Analysis on the DTU Benchmark Dataset
-
Structure-aware Knowledge-guided Heterogeneous Mamba for Zygomaticomaxillary Suture Assessment
-
BadWorld: Adversarial Attacks on World Models
-
PROSE: Training-Free Egocentric Scene Registration with Vision-Language Models
-
Near--Real-Time Conflict-Related Fire Detection in Sudan Using Unsupervised Deep Learning
-
SAMTok: Representing Any Mask with Two Words
-
Clinically Aware Synthetic Image Generation for Concept Coverage in Chest X-ray Models
-
StarOR: Synergizing Tree Search and Test-Time Reinforcement Learning for Optimization Modeling
-
David vs. Goliath in Next Activity Prediction: Argmax vs. LSTM, Transformer, and LLM
-
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation
-
TriAdReview: Triangular Adversarial Review Architecture for Multi-Model Technical Document Generation
-
Beyond the Blood Draw: Explainable Machine Learning for Non-Invasive Dysglycemia Risk Screening
-
Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking
-
Machine learning enables roughness-driven inverse design of milling processes
-
REFLEX: Reflective Evolution from LLM Experience
-
STAR-NT: Spatiotemporal Acceleration of Real-Time Neural Transparency Rendering
-
ACCORD: Action-Conditioned Contextual Grounding for Language Agents
-
LLM Jaggedness Unlocks Scientific Creativity
-
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
-
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
-
Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw
-
Surpassing Scale by Efficiency: A Compact 135M Parameter Foundational LLM Natively Adapted for the Bangla Language
-
How Much Can We Trust LLM Search Agents? Measuring Endorsement Vulnerability to Web Content Manipulation
-
Generative causal testing to bridge data-driven models and scientific theories in language neuroscience
-
Semantic-Preserving Prompt Hijacking: A Black-Box Adversarial Attack on Auto-Prompt Optimization
-
Multiple Descents in Deep Learning as a Sequence of Order-Chaos Transitions in LSTM Networks
-
Double-Helix Vision (DH-V2): A Geometry-Based Visual Sampler for Bandwidth-Constrained Perception
-
CPS4: Class Prompt driven Semi-Supervised Spine Segmentation with Class-specific Consistency Constraint
-
MIRAGE: Runtime Scheduling for Multi-Vector Image Retrieval with Hierarchical Decomposition
-
Is Code Better Than Language for Algorithmic Reasoning
-
Reinforcement Learning for LLM-based Event Forecasting
-
ChatPlanner: A Large Language Model Framework for Personalized Public Transit Routing
-
Machine Learning-Driven Chemical Reactor Network Modeling of the Sandia-D Flame
-
HorusEye: Language as Dynamic Attention for Emergency Visual Analysis
-
SkillVetBench: LLM-as-Judge for Multi-Dimensional Security Risk Evaluation in Open-Source LLM Agent Skills
-
MapDream: Task-Driven Map Learning for Vision-Language Navigation
-
Oops, Wait: Discourse Tokens Matter in Reasoning Model
-
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
-
MixTeX: Data-Efficient LaTeX OCR via Synthetic Pretraining and Limited Fine-Tuning
-
From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents
-
Let Them Steal: Trapping Large Language Model Extraction Attacks with Knowledge Honeypot
-
Tangram: Unlocking Non-Uniform KV Cache Compression for Efficient Multi-turn LLM Serving
-
Graphical-Probabilistic Modeling of Generative Flows in LLM-Native Software Systems
-
Latent Action Pretraining Through World Modeling
-
From ASR to ASP: Evaluating Prompt Attack Vulnerabilities Against Open-Source LLMs
-
The Answer Lies Within: Self-Derived Rewards Enable Explainable Relation Extraction
-
Building Customer Support AI Agents at 100M-User Scale: An Evaluation-Driven Framework
-
FasterPy: An LLM-based Code Execution Efficiency Optimization Framework
-
Beyond Text-to-SQL: An Agentic LLM System for Governed Enterprise Analytics APIs
-
Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation
-
Are Online Skill and Memory Modules Always Worth Their Tokens? A Budget-Constrained Study of Web Agents
-
Snyk VulnBench JS 1.0: Can LLMs Find the Same Bugs Twice?
-
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
-
Understanding Scam Trends and Rail Paths from Reddit Self-Disclosure Narratives
-
Speaking the Language of Science: Toward a General-Purpose Generative Foundation Model for the Natural Sciences
-
RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments
-
LiteOdyssey: A Lightweight Reasoning AI Agent for Interpretable Rare-Disease Diagnosis
-
RAID: Semantic Graph Diffusion for True Cold-Start and Cross-Lingual Forecasting
-
MotionVLA: Vision-Language-Action Model for Humanoid Motion
-
Human genetic evidence is associated with drug approval across therapeutic areas: an observational analysis of 26,278 target-disease pairs with temporal validation and feature ablation
-
AutoDojo: Adaptive Attacks Expose Superficial Defenses and User-Underspecification Limits in LLM Agents
-
Hierarchical Modeling of ICD Codes in EHR Foundation Models
-
Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment
-
Combining Retrieval-Augmented Text Generation with LLMs for Reading Content Recommendations
-
Simplifying the Modeling of Arbitrary Conditionals in Natural Language
-
LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control
-
Phishing Email Detection Using Large Language Models
-
XFlow: An Executable Protocol Programming System for Reliable Multi-Agent Workflows
-
LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
-
Compositional Reasoning Depth Predicts Clinical AI Failure: Empirical Evidence Consistent with Transformer Compositionality Limits in Electronic Health Record Question Answering
-
Recurrent Reasoning on Symbolic Puzzles with Sequence Models
-
SPARK: Security Knowledge Priming and Representation-Guided Knowledge Activation for LLM-based Secure Code Generation
-
JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks
-
Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability
-
GUITrans2Act: Understanding User Operational Behaviors from Mobile GUI Interactions with Vision-Language Models
-
tap: A File-Based Protocol for Heterogeneous LLM Agent Collaboration
-
SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents
-
Crypto x AI, AI x Crypto: A Survey
-
Closing the Reflection Gap: A Free Calibration Bonus for Agentic RL
-
Leave-One-Out-, Bootstrap- and Cross-Conformal Anomaly Detectors
-
Can Deep Neural Networks Improve Compression of Very Large Scientific Data?
-
Learning Urban Access Costs from Origin-Destination Flows via Inverse Optimal Transport
-
Code Correctness Signals in LLM Hidden States: Pre-Generation Probing and Repair Geometry
-
Compressed Computation is (probably) not Computation in Superposition
-
A Statistical and Machine Learning Framework for Operational Threshold Detection and Deployable Dispatch Controller Development in Hydrogen Multi-Energy Systems
-
Recipe-Controlled Decoder Audit for Structural Knowledge-Graph Completion
-
BigPower: Hierarchical Source-Level Module Power Estimation for CPUs with Large Language Models
-
Order Is Not Control: Driven-Dissipative Response Laws Across Artificial and Biological Systems
-
Efficient Rationale-based Retrieval: On-policy Distillation from Generative Rerankers based on JEPA
-
"I Didn't Make the Micro Decisions": Measuring, Inducing, and Exposing Goal-Level AI Contributions in Collaboration
-
Is ChatGPT Fair for Recommendation? Evaluating Fairness in Large Language Model Recommendation
-
MASLab: A Unified and Comprehensive Codebase for LLM-based Multi-Agent Systems
-
Simulating Students' Java Programming Errors with Large Language Models
-
Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP
-
Natively Unlearnable Large Language Models
-
Small LLMs: Pruning vs. Training from Scratch
-
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
-
Beyond the Training Distribution: Evaluating Predictions Under Distribution Shift and Selection Bias
-
Quantizing Time-Series Models As Dynamical Systems: Trajectory-Based Quantization Sensitivity Score
-
Listening with Attention: Entropy-Guided Explainability for Transformer-Based Audio Models
-
Temporal Backtracking Search for Test-time Generative Video Reasoning
-
Avatar V: Scaling Video-Reference Avatar Video Generation
-
From Sorting Algorithms to Scalable Kernels: Bayesian Optimization in High-Dimensional Permutation Spaces
-
Neither Parallel Nor Sequential: How DiffusionGemma Actually Commits Tokens
-
MooMIns -- Monocular 3D Reconstruction and Object Pose Estimation from Multiple Instances
-
SANA: What Matters for QA Agents over Massive Data Lakes?
-
From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails
-
A Fixed-Point Neural Operator for Size- and Functional-Transferable Hamiltonian Prediction
-
TwinBI: An Agentic Digital Twin for Efficient Augmented Interactions with Business Intelligence Dashboards
-
NeST: Neuron Selective Tuning for LLM Safety
-
WAM4D: Fast 4D World Action Model via Spatial Register Tokens
-
FEMOT: Multi-Object Tracking using Frame and Event Cameras
-
I'm Sorry Driver, I'm Afraid I Can't Do That: Appraising the Safety of LLMs within Automotive Contexts
-
The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation
-
AdaSR: Adaptive Streaming Reasoning with Hierarchical Relative Policy Optimization
-
Flex4DHuman: Flexible Multi-view Video Diffusion for 4D Human Reconstruction
-
One Token to Fool LLM-as-a-Judge
-
HairPort: In-context 3D-aware Hair Import and Transfer for Images
-
ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models
-
Two Wrongs, No Right: Auditing Social-Desirability Bias in LLM Annotators for Computational Social Science
-
Recursive Agent Harnesses
-
GEASS: Gated Evidence-Adaptive Selective Caption Trust for Vision-Language Models
-
ARROW: Augmented Replay for RObust World models
-
TokaMark: A Comprehensive Benchmark for MAST Tokamak Plasma Models
-
OCOO-T : A Simple and Scalable Virtual Cell Model for Transcriptional Perturbation Response Prediction
-
LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories
-
DSAEval: Evaluating Data Science Agents on a Wide Range of Real-World Data Science Problems
-
Standardized Methods and Recommendations for Green Federated Learning
-
TAB-PO: Preference Optimization with a Token-Level Adaptive Barrier for Token-Critical Structured Generation
-
Reliability of Probabilistic Emulation of Physical Systems
-
Entropic Mirror Monte Carlo
-
Arbor: Tree Search as a Cognition Layer for Autonomous Agents
-
Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents
-
Prefill Awareness in Large Language Models
-
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
-
GeoDial: A Multimodal Conversational Tutoring Dataset for Geometry Problem-Solving with Visual Tutor Turns
-
HybridCodeAuthorship: A Benchmark Dataset for Line-Level Code Authorship Detection
-
Perceive, Interact, Reason: Building Tool-Augmented Visual Agents for Spatial Reasoning
-
Diffusion Transformer World-Action Model for AV Scene Prediction
-
Multi-Turn Reasoning When Context Arrives in Pieces: Scalable Sharding and Memory-Augmented RL
-
A Quantitative Experimental Repeated Measures Study of Training Dynamics in a Small Llama Style Language Model Under a Compute-Aware Token Budget
-
Democracy in the Era of Artificial Intelligence
-
OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data
-
HalluJudge: A Reference-Free Hallucination Detection for Context Misalignment in Code Review Automation
-
CodeAlchemy: Synthetic Code Rewriting at Scale
-
Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm
-
WebChallenger: A Reliable and Efficient Generalist Web Agent
-
Trust and Reliance on AI in Education: AI Literacy and Need for Cognition as Moderators
-
Mitigating hallucinations in healthcare LLMs with granular fact-checking and domain-specific adaptation
-
HarDBench: A Benchmark for Draft-Based Co-Authoring Jailbreak Attacks for Safe Human-LLM Collaborative Writing
-
Operator Fusion for LLM Inference on the Tensix Architecture
-
CITRAS-FM: Tiny Time Series Foundation Model for Covariate-Informed Zero-Shot Forecasting
-
Exploring the Design Space of Reward Backpropagation for Flow Matching
-
T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains
-
Towards Autonomous Accelerator Design: FPGA Accelerator Generation with SECDA
-
Constructing coherent spatial memory in LLM agents through graph rectification
-
BadRobot: Jailbreaking Embodied LLM Agents in the Physical World
-
Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models
-
How can we assess human-agent interactions? Case studies in software agent design
-
Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement
-
RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning
-
MedFeat: Model-Aware and Explainability-Driven Feature Engineering with LLMs for Clinical Tabular Prediction
-
Do LLMsMakeNeural Distinguishers Wise?
-
Bittensor Agent Arenas as a Trajectory Primitive: Distilling a Shopping Agent from ShoppingBench Subnet Traces
-
Local Is Not a Sufficient Privacy Boundary: Governing OS-Integrated On-Device AI
-
Linguistically Augmented Audio Speech Data (LinguAS)
-
Test-time Adversarial Takeover: A Real-time Hijacking Interface against Robotic Diffusion Policies
-
Learning Evidence Highlighting for Frozen LLMs
-
Using the YOLOv12 Model for Verifying the Correct Color Sequence of Wires in Network Cables (Patch Cords) on the Production Line
-
OpenRTLSet: A Fully Open-Source Dataset for Large Language Model-based Verilog Module Design
-
What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects
-
Who Wrote the Book? Detecting and Attributing LLM Ghostwriters
-
Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents
-
Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
-
LLM-Based Code Documentation Generation and Multi-Judge Evaluation
-
Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans
-
Time Series as Language: A Universal Tokenizer for General-Purpose Time Series Foundation Models
-
CleanPatrick: A Benchmark for Image Data Cleaning
-
Ethical and Technical Limits of Deepfake Speech Datasets
-
ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity
-
Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization
-
WorldOlympiad: Can Your World Model Survive a Triathlon?
-
A Constrained Natural-Language Interface for Variational Multi-Physics Finite Element Simulations in FEniCS
-
Neutrality Bites: Gender Representation in AI-Generated Animal Stories
-
Rosetta Memory: Adaptive Memory for Cross-LLM Agents
-
Evaluating Advanced Prompting on Gemini Flash for Multi-Hop Biomedical QA
-
Multimodal Large Language Models as Synthetic Participants in Video-Based Studies: An Evaluation
-
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
-
BRAIN: Bayesian Reasoning via Active Inference for Agentic and Embodied Intelligence in Mobile Networks
-
Capacity, Not Format: Rethinking Structured Reasoning Failures
-
TheoremBench: Evaluating LLMs on Theorem Proving in Formal Mathematics
-
An Effective Router for Vision-Language Model Selection
-
Beyond Pass Rate: A Multilingual, Execution-Grounded Evaluation of Open Code LLMs
-
The CIFAR Synthetic Evidence Corpus for Detecting AI-Generated Evidence
-
The AI Epistemic Deference Index: A Continuous Measure of Sycophancy
-
Where Instruction Hierarchy Breaks: Diagnosing and Repairing Failures in Reasoning Language Models
-
Are Two Datasets Close Enough With Statistical Significance? A Kernel Distributional Closeness Testing Approach
-
SAGE: An LLM-driven Self Reflective Agentic Framework for Fraud Detection
-
Vector Space of Cycles
-
DIVERGE: Diversity-Enhanced RAG for Open-Ended Information Seeking
-
Distilling Safe LLM Systems via Soft Prompts for On Device Settings
-
Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution
-
GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model
-
Baichuan-M4: A Clinical-Grade Medical Agent System for Continuous Care
-
LLM-Orchestrated Conformance Checking in Stroke Care Without Computer-Interpretable Guidelines
-
Echo-DM: Ultrasound Marker Removal via Conditional Latent Diffusion and Region-Aware Fusion
-
Reconstructing and forecasting disease trajectories of patients with Alzheimer's disease using routine data in resource-constrained settings
-
Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning
-
Reflection in the Dark: Exposing and Escaping the Black Box in Reflective Prompt Optimization
-
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
-
When Do Diffusion Models learn to Generate Multiple Objects?
-
Margin-Adaptive Confidence Ranking for Reliable LLM Judgement
-
Where Does the Answer Come From? Benchmarking View-Level Visual Evidence Identification in Multi-View MLLMs for Autonomous Driving
-
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors?
-
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
-
DeepMine-Mamba: Mitigating Information Dilution in Mamba-Based State Space Models for Document Image Binarization
-
Leveraging NeRF-Rendered Images for 3D Gaussian Splatting
-
X-Palm: Paired Multispectral-to-Smartphone Dataset for Cross-Domain Palmprint Authentication
-
GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization
-
SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing
-
MedVeriSeg: Teaching LISA-Like Medical Segmentation Models to Verify Query Validity Without Extra Training
-
Function-Vector Heads Are Two Populations: Writers and Cancellers in In-Context Learning
-
Automating the Expert Eye: A System-Agnostic Deep Learning Framework for Rare Event Discovery in Imbalanced Force Spectroscopy
-
Learning from flowsheets: A generative transformer model for autocompletion of flowsheets
-
DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation
-
Learning Quantized Continuous Controllers for Integer Hardware
-
XCR-Bench: Benchmarking Cross-Cultural Reasoning in LLMs via Culture-Specific Items and Hall's Triad
-
Implementing Grassroots Logic Programs with Multiagent Transition Systems and AI (Full Version)
-
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
-
CHIMERA-Bench: A Benchmark Dataset for Epitope-Specific Antibody Design
-
DynamicPO: Dynamic Preference Optimization for Recommendation
-
Context Over Compute Human-in-the-Loop Outperforms Iterative Chain-of-Thought Prompting in Interview Answer Quality
-
Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attacks
-
CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models
-
Decoding Naturalistic Emotion Dynamics from the Brain: An LLM-Enhanced Regression Framework
-
De novo molecular generation with optical property preconditioning at the token level
-
AliyunConsoleAgent: Training Web Agents in Real-World Cloud Environments via Distillation and Reinforcement Learning
-
Trajectory Geometry of Transformer Representations Across Layers
-
A Unifying Framework for Concept-Based Representational Similarity
-
Community-Specific Slang and Entity Detection via Semantic Shift in Fine-Tuned Language Models
-
CRANE: Knowledge Editing for Reasoning MLLMs
-
Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models
-
Provably Efficient Personalized Multi-Objective Bandits with Proactive Conversational Queries
-
Agentic Search for Counterfactual Recourse under Fixed LLM Budgets
-
Sample-Efficient LLM-Based Detection of Malicious Web Server Logs with Forensically Explainable Reasoning
-
SafeRun: Enabling Determinism in LLM Planning for Running
-
LargeMonitor: Monitoring Online Task-Free Continual Learning via Large Pretrained Models
-
Self-Paced Curriculum Reinforcement Learning for Autonomous Superbike Racing in Simulation
-
Emergence of Context Characteristics Sensitivity in Large Language Models
-
Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO
-
FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks
-
Language-based Trial and Error Falls Behind in the Era of Experience
-
Executable World Models for ARC-AGI-3 in the Era of Coding Agents
-
Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models
-
$\alpha$-PFN: Fast Entropy Search via In-Context Learning
-
Uncertainty-Aware LLM-Guided Policy Shaping for Sparse-Reward Reinforcement Learning
-
You Only Landmark Once: Lightweight U-Net Face Super Resolution with YOLO-World Landmark Heatmaps
-
The Identity Trap in EEG Foundation Models: A Diagnostic Audit
-
How AI Agents Reshape Knowledge Work: Autonomy, Efficiency, and Scope
-
Agentic Large Language Models for Automated Structural Analysis of 3D Frame Systems
-
What Your Posts Reveal: A Benchmark and Agentic Framework for User-Level Privacy Leakage on Social Media
-
MetaConfigurator: AI-Assisted RDF Authoring from JSON Data
-
DIFFRACT: Neuralized Utility Maximization for Wireless Networks by Differentiable Programming
-
Textual Supervision Enhances Geospatial Representations in Vision-Language Models
-
Exploring Flow-Lenia Universes with a Curiosity-driven AI Scientist: Discovering Diverse Ecosystem Dynamics
-
TSAQA: Time Series Analysis Question And Answering Benchmark
-
SWE-IF: Aligning Code Evaluation with Human Preference
-
SEEK: Steering LLM Reasoning for RAG via Internal Reasoning Sketches
-
Explain Like I'm 5 or Whatever I Choose: Evaluating the Interactive Potential of Language Model Responses
-
Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns
-
Mind the Gap: Bridging Behavioral Silos with LLMs in Multi-Vertical Recommendations
-
Limitations of Normalization in Attention Mechanism
-
Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates
-
Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets
-
DuMate-DeepResearch: An Auditable Multi-Agent System with Recursive Search and Rubric-Grounded Reasoning
-
Topology-Aware Skeleton Detection via Lighthouse-Guided Structured Inference
-
Breaking the Ice: Analyzing Cold Start Latency in vLLM
-
Second-Order Path Kernel Interpolation Formulas in Machine Learning
-
EVA: Evolving Semantic Adversaries for Red-Teaming GUI Agents Against Environmental Injection Attacks
-
Trio: Learning Time-Series Forecasting with Temporal-Spatial-Sample Attention and Structural Causal Priors
-
PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration
-
A Comprehensive Anatomy of Human and DeepSeek-R1 LLM Mathematical Reasoning
-
More Capable, Less Cooperative? When LLMs Fail At Zero-Cost Collaboration
-
Automatic Causal Fairness Analysis with LLM-Generated Reporting
-
LLM-Guided Evolution for Medical Decision Pipelines