Entities · Models
alphaXiv
29 articles tagged with this entity.
-
The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents
-
When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning
-
When do prophets profit in prediction markets?
-
RustMizan: A Compilable, Contamination-Aware Benchmarking Framework for Rust Vulnerabilities
-
SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation
-
Road-Aware Anomaly Segmentation with Query-Guided Polygons and CLIP in Autonomous Driving
-
Collaborative Disagreement Resolution for Scalable Oversight
-
Large Databases Need Small, Open-Weight Language Models
-
Use What You Know: Causal Foundation Models with Partial Graphs
-
PhoneBuddy: Training Open Models for Agentic Phone Use
-
Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints
-
AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning
-
BenchX: Benchmarking AI Models for Cancer Detection and Localization with Demographic and Protocol Biases
-
CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression
-
Metis: Bridging Text and Code Memory for Self-Evolving Agents
-
Human-like autonomy emerges from self-play and a pinch of human data
-
Emergent Alignment
-
Attribute Inference from Interactive Targeted Ads
-
Computational Safety for Generative AI: A Hypothesis Testing Perspective
-
TERMS-Bench: Diagnosing LLM Negotiation Agents Beyond Deal Rate
-
Vernier: Probing Representational Misalignment Behind Lexical Gaps in Causal Reasoning
-
Automatic identification of diagnosis from hospital discharge letters via weakly supervised Natural Language Processing
-
Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models
-
LATTEArena: An Evaluation Framework for LLM-powered Tabular Feature Engineering (Extended Version)
-
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping
-
Taming Perception Jitter: Uncertainty-Aware LiDAR Object Detection for Reliable Motion Classification
-
SoK: Reconstruction Attacks on Synthetic Tabular Data (Insights from Winning the NIST CRC)
-
Rewrite to Translate, Translate to Reward: Reinforcement Learning for Source Rewriting in Machine Translation
-
Depth over Fidelity in Fixed-Budget Noisy Evolution Strategies