Entities · People
arXivLabs
37 articles tagged with this entity.
-
A Unified Detection Framework for AI-Related Content and Artifacts
-
Andha-Dhun: A First Look at Audio Descriptions in Hindi
-
UCSC NLP at SemEval-2026 Task 10: Boundary-Aware Span Extraction and RoBERTa Classification for Conspiracy Detection
-
VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection
-
Evaluating Time Series Foundation Models for Electricity Price Forecasting: Contamination Risk, Distributional Shifts, and Covariate Dependence
-
Governing Generative AI Across Financial Institutions: An SR 26-2-Compatible Framework for Generative AI Risk Control
-
Rethinking Complexity Metrics for LLM-Integrated Applications: Beyond Source Code
-
LLVM-Bench: Benchmarking and Advancing Large Language Models for LLVM Compiler Issue Resolution
-
AgentBound: Verifiable Behavioral Governance for Autonomous AI Agents
-
On the Nonlinearity of Learning Rate Scaling for LLM Training
-
Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing
-
Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries
-
Vision-Language-Action Models: Experimental Insights from a Real-World UR5 Platform
-
Pigeonholing: Bad prompts hurt models to collapse and make mistakes
-
Self-Preference Is Weak or Absent in Verifiable Instruction-Following Revision: A Four-Model Test Under Genuine Authorship
-
Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts
-
SCOPE-FL: A Strategy-proof Chain-based Optimal pareto efficient Federated Learning System
-
Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
-
Toward Accessible Psychotherapy Training Using AI-Driven Interactive Patient Avatars
-
ChatPlanner: A Large Language Model Framework for Personalized Public Transit Routing
-
Let Them Steal: Trapping Large Language Model Extraction Attacks with Knowledge Honeypot
-
Understanding Scam Trends and Rail Paths from Reddit Self-Disclosure Narratives
-
RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments
-
AutoDojo: Adaptive Attacks Expose Superficial Defenses and User-Underspecification Limits in LLM Agents
-
Compositional Reasoning Depth Predicts Clinical AI Failure: Empirical Evidence Consistent with Transformer Compositionality Limits in Electronic Health Record Question Answering
-
Does the Judge Prefer English? Evaluating Language-Switching Invariance in LLM-as-a-Judge
-
TwinBI: An Agentic Digital Twin for Efficient Augmented Interactions with Business Intelligence Dashboards
-
The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation
-
Learning the Universe: Posterior Reliability of Neural Generative Models in High-Dimensional Field-Level Inference of Cosmic Initial Conditions
-
Do LLMsMakeNeural Distinguishers Wise?
-
Multimodal Large Language Models as Synthetic Participants in Video-Based Studies: An Evaluation
-
Vector Space of Cycles
-
Where Does the Answer Come From? Benchmarking View-Level Visual Evidence Identification in Multi-View MLLMs for Autonomous Driving
-
X-Palm: Paired Multispectral-to-Smartphone Dataset for Cross-Domain Palmprint Authentication
-
Self-Paced Curriculum Reinforcement Learning for Autonomous Superbike Racing in Simulation
-
Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models
-
Mind the Gap: Bridging Behavioral Silos with LLMs in Multi-Vertical Recommendations