Entities · Locations
cs.AI
28 articles tagged with this entity.
-
Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety
-
The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents
-
AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
-
QANTIS: Hardware-Calibrated Sequential POMDP Belief Updates on IBM Heron
-
Danus: Orchestrating Mathematical Reasoning Agents with Fact-Graph Memory
-
AutoMem: Automated Learning of Memory as a Cognitive Skill
-
Modality-Driven Search with Holistic Trace Judging for ARC-AGI-2
-
Selective Memory Retention for Long-Horizon LLM Agents
-
Evidence-Informed LLM Beliefs for Continual Scientific Discovery
-
How Much Due Diligence Before You Bid? Learning in Intractable Takeover Auctions
-
MARS: A neurosymbolic approach for interpretable drug discovery
-
When Summaries Distort Decisions: Information Fidelity in LLM-Compressed Financial Analysis
-
DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forums
-
Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems
-
Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System
-
A Multi-Agent system for Multi-Objective constrained optimization
-
PCBSchemaGen: Reward-Guided LLM Code Synthesis for Printed Circuit Boards (PCB) Schematic Design with Structured Verification
-
ForecastBench-Sim: A Simulated-World Forecasting Benchmark
-
In-Context Environments Induce Evaluation-Awareness in Language Models
-
Towards Next-Generation Healthcare: A Survey of Medical Embodied AI for Perception, Decision-Making, and Action
-
Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations
-
Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact
-
Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem Proving
-
Orchestra-o1: Omnimodal Agent Orchestration
-
When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models
-
A History-Aware Visually Grounded Critic for Computer Use Agents
-
TQA-Bench: Evaluating LLMs for Multi-Table Question Answering
-
Playing Devil's Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy