Entities · Labs
GotitPub
505 articles tagged with this entity.
-
InductWave: Inductive Multi-Hop Logical Query Answering on Knowledge Graphs
-
Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks
-
Latency-Constrained DNN Architecture Learning for Edge Systems using Zerorized Batch Normalization
-
A knowledge-augmented dataset of high-risk driving scenarios with LLM annotations for autonomous driving
-
Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity
-
HPR-SAM: Hierarchical Probabilistic Representation Learning for Prompt-free SAM-based Medical Image Segmentation
-
SynthAVE: Scalable Synthetic Labeling for E-Commerce with LLM-Arena Validation
-
tsbootstrap: Distribution-Free Uncertainty Quantification and Conformal Prediction for Time Series
-
Health System Scale Semantic Search Across Unstructured Clinical Notes
-
A Good Initialization is All You Need for Faithful Visual Attribution
-
STAGformer: A Spatio-temporal Agent Graph Transformer for Micro Mobility Demand Forecasting
-
Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild
-
MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models
-
PALS: Percentile-Aware Layerwise Sparsity for LLM Pruning
-
Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data
-
UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation
-
x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability
-
MambaGaze: Bidirectional Mamba with Explicit Missing Data Modeling for Cognitive Load Assessment from Eye-Gaze Tracking Data
-
Why does Deep Learning Improve Visual SLAM?
-
RMISC: A Large-scale Real-world Multivariate Corpus for Time Series Foundation Models
-
Lean-Quantum: Toward AI-Assisted Formalization of Quantum Information
-
Umm... With Transformers? Insights from Filled Pause Use across Four Slavic Parliaments
-
From Blueprint to Reality: Modeling and Applying Putnam's Social Capital Theory with LLM-based Multi-agent Simulations
-
CurateEvo: Data-Curation Evolving for Agentic Post-Training
-
ResonatorLM: Causal Resonant Field Mixing for Efficient Long-Context Language Modelin
-
CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration
-
MobileWan: Closing the Quality Gap for Mobile Video Diffusion
-
AlayaWorld: Long-Horizon and Playable Video World Generation
-
Parameter-Free Encoders Remain Viable for RDB Foundation Models
-
Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development
-
PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages
-
Separating Representation from Reconstruction Enables Scalable Text Encoders
-
Rethinking Depth Pruning for Vision Transformers: A Heterogeneity-Aware Perspective
-
rePIRL: Learn PRM with Inverse RL for LLM Reasoning
-
VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning
-
Rethinking Scientific Discovery in an Agentic Era
-
MPSelectTune: Prompt-type Selection for Fine-tuning improves Concept Unlearning in LLMs
-
Is Agentic Code Review Helpful? Mining Developers' Feedback to CodeRabbit Reviews in the Wild
-
A non-invasive video-based method for individual identification of wildlife using gait dynamics
-
EVA-Client: A Unified Data Collection, Inference, and Deployment Framework for Embodied Policies on Real Robots
-
IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
-
Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation
-
FuseMamba-VD: Dual Branch VideoMamba with Gated Class Token Fusion for Violence Detection
-
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment
-
msPCA: An R Package for Sparse PCA with Multiple Components
-
QuantFlow: A Federated Mamba-Based Post-Transformer Foundation Model for Time-Series Forecasting
-
Spectral Rewiring for Exploration, Purification, and Model Merging
-
PhenoNEST: A Neuro-Symbolic Framework for Ontology-Aware Multimodal Plant Phenotyping and Trait Discovery
-
WeightCLIP: Aligning Datasets and Models for Weight Space Learning
-
Adversarial LassoNet: Robust Feature Selection via Stability-Driven Sparse Learning
-
Fusion: A Framework for Unified Sequential Token AdaptatIon in VisiOn TraNsformers
-
RABBiT: Rapidly adaptive BOLD foundation model via brain-tuning for accurate zero-shot and few-shot prediction of speech-elicited responses in the brain
-
Exploring SAM Supervision for Fine-Grained UAV Target Segmentation under Data Scarcity
-
RayTun3R: Online Camera Adaptation in 3D Foundation Models
-
RotateAttention: RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation
-
FORGE: Research-Trajectory Hijacking Attacks on Deep Research Agents
-
OmniDS: Dual-Stream Context Fusion for Omnidirectional Depth from Fisheye Cameras
-
The Classics at SemEval-2026 Task 3: Combining Transformer Models and LLM-Generated Annotations for Dimensional Aspect-Based Sentiment Analysis
-
Decentralized Aggregation of LLM Predictions via Wagering Mechanisms
-
Angry but Accurate: Detecting and Profiling the Counter-Misinformation Ecosystem on Twitter
-
QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs
-
Global Logic and Local Search: Dual-Stream Multimodal In-Context Learning for Verifiable Industrial Anomaly Detection
-
You Frame It: How Conceptual Representations Shape LLM Detection and Reasoning about Antisemitism
-
CertMix: Certified, Data-Efficient Metamaterial Design by Affine Mixing of Aligned Neural-Implicit Weight Spaces
-
Grokking Is Conditional and Fragile: A Fully-Tractable, Multi-Seed Study at 12K Parameters
-
Localized LoRA-MoE: Block-wise Low-Rank Experts With Adaptive Routing
-
Incentivizing Vision Language Models to Search for Long Video Question Answering
-
MTEB-PT: A Text Embedding Benchmark for Brazilian Portuguese
-
Language Models Represent and Transform Concepts with Shared Geometry
-
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception
-
FairFlow: Demystifying and Mitigating Stereotype Bias in Text-to-Image Diffusion Transformers
-
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
-
DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation
-
ReCal3R: Reliability-Calibrated Learning Rates for Streaming 3D Reconstruction
-
SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass
-
DynaWM: A Base-VLA-Guided World Foundation Model for Moving-Object Manipulation
-
Risk-Constrained Freshness-Aware Semantic Caching for Open-Web Retrieval-Augmented LLMs
-
The Remarkable Effectiveness of Providing AI Agents with Natural Language Tools: A Replication Study Validating NLT Performance Across 14 Models
-
Framework of Thoughts: A Foundation Framework for Dynamic and Optimized Reasoning based on Chains, Trees, and Graphs
-
Don't Make Models Guess Security and Safety: Symbolic Guardrails for Domain-Specific AI Agents
-
Legible-by-Construction: Attention and End-to-End Transformers
-
Do ECG Foundation Models Transfer to Rare Cardiac Diseases? Evidence from Brugada Syndrome Detection
-
MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution
-
Conversational Human Audio-visual Talking Dialogue Generation
-
VISTA: Auditing Semantic Divergence in Vision-Language Models
-
How Do Diffusion Classifiers Decide? A Bias-Centric Evaluation
-
OmniLayout: A Schematic-Coupled Multimodal Benchmark for Constraint-Aware Geometric Reasoning in PCB Layout
-
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
-
Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling
-
Anti-Prompt: Image Protection against Text-Guided Image-to-Video Generation
-
Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling
-
UA-ChatDev: Uncertainty-Aware Multi-Agent Collaboration for Reliable Software Development
-
Aria: An Agent For Retrieval and Iterative Auto-Formalization via Dependency Graph
-
Zeus: Towards Tuning-Free Foundation Model for Time Series Analysis
-
LLM-Empowered Multimodal Fusion Framework for Autonomous Driving: Semantic Enhancement and Channel-Adaptive Design
-
WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs
-
From Lab to Reality: A Practical Evaluation of Deep Learning Models and LLMs for Vulnerability Detection
-
Generic Expert Coverage for Pruning SparseMixture-of-Experts Language Models
-
kNNGuard: Turning LLM Hidden Activations into a Training-Free Configurable Guardrail
-
EPnG: Adaptive Expert Prune-and-Grow for Parameter-Efficient MoE Fine-tuning
-
Single-Channel EEG-Based Cognitive Load Assessment in Online Learning: A Hybrid Deep Learning Approach
-
Object Aligner: A Configurable JSON Schema Similarity Score for Graphs, Applied to LLM Prompt Optimization
-
Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias
-
From Experiments to Expertise: Scientific Knowledge Consolidation for AI-Driven Computational Physics
-
Know Your Source: A Public Knowledge Store for Media Background Checks
-
DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation
-
DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing
-
CreativityNeuro: Steering Language Model Weights to Improve Divergent Thinking and Reduce Mode Collapse
-
OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration
-
Safe and Adaptive Cloud Healing: Verifying LLM-Generated Recovery Plans with a Neural-Symbolic World Model
-
Autonomous discovery of traffic laws with AI traffic scientists
-
Atomic Task Graph: A Unified Framework for Agentic Planning and Execution
-
Safeguarding LLM Agents from Misalignment through Provenance Analysis
-
Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction
-
ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning
-
Scaling Trends for Lie Detector Oversight in Preference Learning
-
DL-VINS-Factory: A Modular Framework for Learned Visual Front-Ends in Visual-Inertial SLAM
-
Towards Metric-Agnostic Trajectory Forecasting
-
"Don't Say It!": Constraints, Compliance, and Communication when Language Models Play Taboo
-
Seahorse: A Unified Benchmarking Framework for Spatiotemporal Event Modeling
-
PRISM: Prioritized Channel Importance with Semi-supervised Domain Adaptation for Cross-Subject EEG Emotion Recognition
-
When Less is More: 8-bit Quantization Improves Continual Learning in Large Language Models
-
LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models
-
SenseWalk: Agent-Based Semantic Trajectory Simulation Powered by Large Language Models in Zoned Environments
-
YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese
-
NoPA: Non-Parametric Online 3D Scene Graph Generation
-
EquiSteer: Cross-Attention Steering Towards a Fairer Text-Guided Image Generation
-
Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments
-
Agentic generation of verifiable rules for deterministic, self-expanding reaction classification
-
Diffusion-GR2: Diffusion Generative Reasoning Re-ranker
-
OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale
-
Beyond Activation Alignment:The Alignment-Diversity Tradeoff in Task-Aware LLM Quantization
-
Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation
-
Persona Without Substrate: Regime-Dependence and the LLM Individuation Problem
-
SemiScope: Disentangling Classifier Tuning and Joint Optimization in Semi-Supervised Security Classification
-
LeNEPA: No-Augmentation Next-Latent Prediction for Time-Series Representation Learning
-
Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training
-
FusionFactory: Fusing LLM Capabilities with Multi-LLM Log Data
-
AGC-Bench: Measuring Artificial General Creativity
-
Concept Alignment Contrast and Long-Short Prompt Memory for Test-Time Adaptation of SAM3 in Medical Image Segmentation
-
From Technical Metrics to User Perception: A User Study of a Multimodal Human-Robot Interaction System for Object Detection and Grasping
-
Robust Text Watermarking for Large Language Models via Dual Semantic Embeddings
-
LOPA: Enhancing Spoken Language Assessment via Latent Ordinal Prototype Alignment
-
Cross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on Romanian
-
SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation
-
Seeing Through Multiple Views: Parameter-Efficient Fine-Tuning via Selective Neurons for Consistent Radiology Report Generation
-
ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models
-
FARS: A Fully Automated Research System Deployed at Scale
-
WIDER-FAIR: An Annotated Version of the WIDER-FACE Dataset for Fairness Evaluation
-
CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning
-
Can LLMs Imagine Moral Alternatives Beyond Binary Dilemmas?
-
Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA
-
AgRefactor: Self-Evolving Agentic Workflow for HLS Compatibility and Performance
-
MultiUAV-Plat: An LLM-Oriented Platform, Benchmark and Framework for Multi-UAV Collaborative Task Planning
-
ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents
-
Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law
-
Horseshoe Priors for Spatial Small Area Estimation: Regular Variation, Tail Robustness, and Deep Learning
-
Fora: From Weight-Space to Function-Space Protection in Capability-Preserving Fine-Tuning
-
A Systematic Approach to Multi-Agent AI from Advanced Regulatory Control Theory: Safe and Auditable LLM Operator Agents for Process Control
-
MOSAIC: Orchestrating Collaborative Knowledge Tracing with Hierarchical Semantic Alignment
-
The Contagion Tensor: A Framework for Measuring Output-Distribution Coupling in Multi-Agent LLM Systems -- and Auditing the Claims It Enables
-
Multimodal Mathematical Reasoning with Diverse Solving Perspective
-
Attractor States Emerge in Multi-Turn LLM Conversations
-
You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
-
Diagnosing and Mitigating Context Rot in Long-horizon Search
-
GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots
-
A Diagnostic Framework and Multi-Evaluator Audit of Evaluator-Driven Preference Dynamics in Self-Adapting LLM Agents
-
PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM Agents
-
DataComp-VLM: Improved Open Datasets for Vision-Language Models
-
Reliability-Prioritized Fine-Grained Generation in Multimodal Large
-
HEARTS: Benchmarking LLM Reasoning on Health Time Series
-
LLM-Guided Planning for Multi-hop Reasoning over Multimodal Nuclear Regulatory Documents
-
A Knowledge Theory of Capital:The Value of Natural and Artificial Intelligence, Volume 1
-
The strength of clinical evidence is recoverable from language model representations but not from their stated grades
-
TriageRA-CCF: Source-Side Clinical Confidence and Coverage Signals for Adaptive Rank Budgeting in Medical LLMs
-
Can LLMs Hire Fairly? Racial Bias in Resume Screening
-
Timesteps of Mamba Align with Human Reading Times
-
SICAGE: Speaker-Independent Culture-Aware Gesture Generation using TED4C-L Dataset
-
Value-Action Alignment in Large Language Models under Privacy-Prosocial Conflict
-
Blackknife: Hard-Label Query-Limited Black-Box Attacks on Heterogeneous Graph Neural Networks
-
CAREBench: A Child-Safety Risk Benchmark for Language Models
-
Beyond IID: How General Are Tabular Foundation Models, Really?
-
AlgoSkill: Learning to Design Algorithms by Scheduling Human-Like Skills
-
Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation
-
SemJoin: Semantic Join Optimization
-
How much of an LLM-generated clinical corpus is actually new? A production-scale measurement of content redundancy for provenance classification
-
Mandol: An Agglomerative Agent Memory System for Long-Term Conversations
-
GLACIER: Rethinking Mass Spectrum Prediction as an Object Detection Problem
-
Towards Evaluating Data Priors for Tabular Foundation Models
-
Legal Domain Adaptation of Modern BERT Models
-
Self-Evolving Agentic Image Restoration via Deliberate Planning and Intuitive Execution
-
An AI agent for treatment reasoning over a biomedical tool universe
-
Agentic Abstention: Do Agents Know When to Stop Instead of Act?
-
SFBench: The SciFy Scientific Feasibility Benchmark
-
Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents
-
Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics
-
How Do LLMs Cite? A Mechanistic Interpretation of Attribution in Retrieval-Augmented Generation
-
Low-cost concept-based localized explanations: How far can we get with training-free approaches?
-
Fuzzing Large Language Models to Elicit Hidden Behaviours
-
One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining
-
RADIANT-PET: Reasoning-Augmented PET/CT Lesion Segmentation with Large Language Models and Reinforcement Learning
-
Dynamic Parsing and Updating Natural Language Specification using VLMs for Robust Vision-Language Tracking
-
SCARCE: Scalable Cascade Analysis for Rare-event Characterisation via Embeddings
-
PromptGNN-sim: Deep Fusion and Alignment of GNN and LLMs for Text-Attributed Graph Learning
-
Online Data Selection for Instruction Tuning via Gaussian Processes
-
Diffusion Fine-tuning with Rewarded Moment Matching Distillation
-
TraceLab: Characterizing Coding Agent Workloads for LLM Serving
-
FlipGuard: Defending Large Language Models Against Quantization-Conditioned Backdoor Attacks
-
Democratic ICAI: Debating Our Way to Steering Principles from Preferences
-
Quantum Generative Diffusion Model for Real-World Time Series
-
RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning
-
Enhancing Numerical Prediction in LLMs via Smooth MMD Alignment
-
A Tree-of-Thoughts Inspired Hybrid Approach for Legal Case Judgement Summarization using LLMs
-
Cross-Platform Chinese Offensive Comment Detection via Dual-Threshold Hard Example Mining
-
VASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned Anchoring
-
Flexformer: Flexible Linear Transformer with Learnable Attention Kernel
-
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context
-
CascadeOcc: Rethinking 3D Occupancy World Models with Cascaded VQ Representations
-
The Context-Ready Transformer
-
From Signals to Transfer: A Factorised Study of Probe-Based Uncertainty Estimation in Large Language Models
-
Otter Weather: Skillful and Computationally Efficient Medium-Range Weather Forecasting
-
When Role-playing, Do Models Believe What They Say?
-
Escaping Iterative Parameter-Space Noise: Differentially Private Learning with a Hypernetwork
-
Where Larger Models Excel: The Primacy of Constraint-Guided Reasoning
-
Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training
-
LAMP: Lane-Aligned Motion Primitives for Feasible Trajectory Prediction
-
Disco-LoRA: Disentangled Composition of Content, Style, and Motion for Multi-concept Video Customization
-
What Do Deepfake Benchmarks Measure? An Audit Using Frozen Self-Supervised Representations
-
TraMP-LLaMA: Generative Interpretability with Decoupled Instruction Tuning for Facial Expression Quality Assessment
-
OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents
-
Know2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models
-
Joint Learning of Experiential Rules and Policies for Large Language Model Agents
-
ShareLock: A Stealthy Multi-Tool Threshold Poisoning Attack Against MCP
-
TEMPO-Diffusion: Temporally Exposed Malicious Poisoning of Diffusion Models
-
TaskNPoint: How to Teach Your Humanoid to Hit a Backhand in Minutes
-
Agentic Analysis for Agentic Infrastructure: An LLM-Powered Pipeline for Comparative Governance of DAO and Corporate AI Protocols
-
COrigami: An AI Pipeline for Co-Designing Flat-Foldable Visually Recognisable Origami
-
Zero-shot Tweet-Level Stance Detection Enhanced by External Knowledge and Reflective Chain-of-Thought Reasoning
-
Heterogeneous Neural Predictivity from Language Models During Naturalistic Comprehension
-
Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets
-
Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding
-
EPTS: Elastic Post-Training Sparsity for Efficient Large Language Model Compression
-
MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources
-
Distill on a Diet: Efficient Knowledge Distillation via Learnable Data Pruning
-
SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion
-
TokenMinds: Pretrained User Tokens and Embeddings for User Understanding in Large Recommender Systems
-
C3-Bench: A Context-Aware Change Captioning Benchmark
-
EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis
-
SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery
-
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents
-
SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models
-
Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors
-
Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints
-
Business as Rulesual: A Benchmark and Framework for Business Rule Flow Modeling with LLMs
-
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?
-
Layer-wise Probing of wav2vec 2.0 and Whisper for Consonant Cluster Reduction in African American English
-
DiffusionBench: On Holistic Evaluation of Diffusion Transformers
-
SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks
-
MM-TRELLIS: Point-Cloud Guided Multi-Modal 3D Vehicle Generation in Autonomous Driving
-
Latent Visual States for Efficient Multimodal Reasoning
-
HANCLIP: A Family of Hyperbolic Angular Negation Vision Language Models
-
Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models
-
DramaDirector: Geometry-Guided Short Drama Generation
-
Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations
-
Are We Ready For An Agent-Native Memory System?
-
RAVEN: A Regime-Aware Variable-context Expert Network for Financial Time Series Forecasting
-
Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning
-
ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection
-
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference
-
BluTrain: A C++/CUDA Framework for AI Systems
-
PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models
-
VeryTrace: Verifying Reasoning Traces through Compilable Formalism and Structured Verification
-
Aligning Audio Captions with Human Preferences
-
Qwen-AgentWorld: Language World Models for General Agents
-
NoLimits.jl: Flexible and Composable Nonlinear Mixed-Effects Modeling in Julia
-
Scaling Laws for Task-Specific LLM Distillation
-
CALIBER: Calibrating Confidence Before and After Reasoning in Language Models
-
NRITYAM: Language Models Meet Art and Heritage of Dance
-
PsyScore: A Psychometrically-Aware Framework for Trait-Adaptive Essay Scoring and ZPD-Scaffolded Feedback
-
BAFIS: Dataset + Framework to assess occupational Bias and Human Preference in modern Text-to-image Models
-
CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs
-
ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference
-
IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
-
ProMUSE: Progressive Multi-modal Uncertainty-guided Staged Evidential Alzheimer Disease Classification
-
Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement
-
Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA
-
The Hidden Evolution of Disguised Visual Context inside the VLM
-
Spectral Retrieval-Augmented Time-Series Forecasting
-
Insulin4RL: Real-Time Insulin Management in the Intensive Care Unit for Offline Reinforcement Learning
-
FloatDoor: Platform-Triggered Backdoors in LLMs
-
Pixel-Level Residual Diffusion Transformer: Scalable 3D CT Volume Generation
-
LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment
-
FLiP: Towards understanding and interpreting multimodal multilingual sentence embeddings
-
A Unified Framework for Efficient Remote Sensing Visual Question Answering: Adapting Dual, Hybrid, and Encoder-Decoder Architectures
-
Beyond Scalar Scores: Exploring LLM-based Metrics for Clinical Significance Evaluation in Radiology Reports
-
Leadership as Coordination Control: Behavioral Signatures and the Recovery-Advantage Boundary in Multi-Agent LLM Teams
-
Improve Large Language Model Systems with User Logs
-
InfoPO: Information-Driven Policy Optimization for User-Centric Agents
-
Enhancing Multilingual Reasoning via Steerable Model Merging
-
Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization
-
EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts
-
OrthoReg: Orthogonal Regularization for Hybrid Symbolic-Neural Dynamical Systems
-
CAOA -- Completion-Assisted Object-CAD Alignment
-
Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States
-
Confidence is Not Reliability: Rethinking MC Dropout in Brain Tumour Segmentation
-
See First, Answer Later: Visual Evidence Pre-Alignment via Sufficiency-Driven RL
-
When Rules Learn: A Self-Evolving Agent for Legal Case Retrieval
-
Dissecting model behavior through agent trajectories
-
DecoSearch: Complexity-Aware Routing and Plan-Level Repair for Text-to-SQL
-
Feynman Kac Reweighted Schr\"odinger Bridge Matching for Surface-Based Tau PET Harmonization
-
Comprehensive pKa Data Augmentation from Limited Real Data through an Engineered Models-Quantum Framework
-
RooseBERT: A New Deal For Political Language Modelling
-
MoSE: Mixture of Slimmable Experts for Efficient and Adaptive Language Models
-
RadSEM: A Finding-by-Finding Metric for Clinical Consistency in Radiology Reports
-
MedicalAgentsBench for Complex Medical Reasoning: Comparing Internalized Reasoning Models versus Externalized Agent-based Frameworks
-
Deep Reinforcement Learning for Minimum Zero-Forcing Sets
-
AnchorKV: Safety-Aware KV Cache Compression via Soft Penalty with a Refusal Anchor
-
LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams
-
S4oP: Operator-level Pruning of Structured State Space Models for Resource-Constrained Devices
-
Blueprint First, Model Second: A Framework for Deterministic LLM Workflow
-
TaFD: Threat-Aware Frequency Decoupling for Adversarial Robustness against Heterogeneous Attacks
-
EventDrive: Event Cameras for Vision-Language Driving Intelligence
-
Advances in 4D Representation: Geometry, Motion, and Interaction
-
SCC-Loc: A Unified Semantic Cascade Consensus Framework for UAV Thermal Geo-Localization
-
LLMs Infer Cultural Context but Fail to Apply It When Responding
-
NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama
-
Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing
-
Nothing from Something: Can a Language Model Discover 0?
-
Vision-language models for chest radiography do not always need the image
-
Structural Role Injection in Handlebars-Templated LLM Prompts: Triple-Brace Interpolation, Delimiter Family, and the Limits of HTML Auto-Escaping
-
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack
-
ParkingTransformer: LLM-Enhanced End-to-End Trajectory Planning for Autonomous Parking
-
HRDX: A Large-Scale Vector HD-Map Dataset
-
Mask-Proof: An LLM-based Automated Data Curation Pipeline on Mathematical Proofs
-
CrossMaps: Confidence-Aware Open-Vocabulary Semantic Mapping for Rover Navigation
-
On the Adversarial Robustness of Multimodal LLM Judges
-
GRACE: Boosting Video MLLMs with Grounded Action-Centric Evidence for Viewer Sentiment Prediction
-
MVEB: Massive Video Embedding Benchmark
-
Constitutional Value Potentials: reading and steering internal priority margins in language models
-
Beyond the Blood Draw: Explainable Machine Learning for Non-Invasive Dysglycemia Risk Screening
-
Pre-Training for Simulation-Based Science: A Study on Jet Foundation Model Training Objectives
-
Distilling latent electrostatics from foundation machine learning interatomic potentials
-
ACC: Compiling Agent Trajectories for Long-Context Training
-
The Value Axis: Language Models Encode Whether They're on the Right Track
-
FlowMPC: Improving Flow Matching policies with World Models
-
LLM Judges Have Dark Current: A Psychometric Datasheet for LLM-as-a-Judge Evaluation
-
The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model Lineages
-
Pepti-Agent: An AI Agent for Peptide Design and Optimization
-
Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability
-
LatentGym: A Testbed For Cross-Task Experiential Learning With Controllable Latent Structure
-
Do LLMs Reliably Identify Correct Information Units in Aphasic Discourse?
-
GAS-Leak-LLM: Genetic Algorithm-Based Suffix Optimization for Black-Box LLM Jailbreaking
-
EdgeZSAD: Practical Zero-Shot Anomaly Detection on Edge Devices
-
MIRAGE: Auditing Anti-Muslim Bias in Frontier LLMs Across Reasoning, Agentic, and Time-Coupled Conditions
-
DoubtProbe: Black-Box Jailbreak Defense via Structural Verification and Semantic Auditing
-
Beyond Layer Importance in Layer-wise Sparsity: An Inter-Layer Perturbation-Absorption Perspective
-
BALTO: Balanced Token-Level Policy Optimization for Hallucination Mitigation
-
Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs
-
MosaicQuant: Inlier-Outlier Disaggregation for Unified 4-Bit LLM Quantization
-
Sub-Semantic Image Segmentation
-
Scaling LLM Reasoning from Minimal Labels: A Semi-Supervised Framework with a Lightweight Verifier
-
Few-Shot Biomedical Relation Extraction with Large Language Models: A Viable Alternative to Supervised Learning?
-
PACUTE: Phonology-, Affix-, and Character-level Understanding of Tokens for Filipino
-
LLM-Assisted Stance Detection in Scientific Discourse: A Test Case in Bayesian Cognitive Science
-
Long-Context Modeling via GSS-Transformer Hybrid Architecture with Learnable Mixing
-
UXBench: Measuring the Actionability of LLM-Generated UX Critiques
-
AgentLeak: A Benchmark for Internal-Channel Privacy Leakage in Multi-Agent LLM Systems
-
Naive Visual Memory is Not Enough: A Failure-Mode Study of GUI Agents
-
FLaRA: Predicting Future Latent Representations for Accident Anticipation
-
VideoWeave: Unlocking Geometric Consistency in Video Generation via Joint Geometry-Video Modeling
-
RT-VLA: Real-Time Vision-Language-Action Models via Knowledge Distillation
-
The Culture Funnel: You Can't Align What isn't in the Data
-
EiCAP: Beyond Fluency, Probing and Improving Emotional Intelligence in LLMs via Psychologically Grounded Multi-Turn Dialogue
-
Hybrid Classical-Quantum Variational Autoencoder for Neural Topic Modeling
-
Pix2Fact: When Vision Is Not Enough -- Benchmarking Fine-Grained VQA with Web Verification on High-Resolution Real-World Scenes
-
A General Framework for Decision Trees via Bregman Divergences
-
Poker Arena: Multi-Axis Profiling of Strategic Reasoning and Memory in LLMs
-
ADORE: Iterative Query Expansion with Retrieval-Grounded Relevance Feedback
-
Realizing Native INT8 Compute for Diffusion Transformers on Consumer GPUs: A Fused INT8 GEMM Kernel for Ideogram 4.0
-
Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH
-
AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models
-
Low-Burden LLM-Based Preference Learning: Personalizing Assistive Robots from Natural Language Feedback for Users with Paralysis
-
Fusing Stylometric and Embedding Systems to Estimate Authorship Likelihood Ratios in Japanese
-
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
-
Toward 360-Degree Indoor Panorama Editing via Tuning-Free Diffusion Model with Refocusing Cross-Attention
-
Beyond Perplexity: UTF-8 Validity in Byte-aware Language Models
-
SciDef: Datasets and Tools for Automated Definition Extraction from Scientific Literature with LLMs
-
A Fixed-Point Neural Operator for Size- and Functional-Transferable Hamiltonian Prediction
-
Revisiting Vehicle Color Recognition in Long-Tailed Surveillance Scenarios
-
Causal Inference with Generative Artificial Intelligence: Application to Texts as Treatments
-
Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review
-
Small LLMs for Biomedical Claim Verification: Cost-Effective Fine-Tuning, Structural Dataset Shortcuts, and Cross-Domain Generalization
-
Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents
-
SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation
-
When Smaller Wins: Dual-Stage Distillation and Pareto-Guided Compression of Liquid Neural Networks for Edge Battery Prognostics
-
When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering
-
EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
-
Benchmarking AI Agents for Addressing Scientific Challenges Across Scales
-
ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs
-
MLUBench: A Benchmark for Lifelong Unlearning Evaluation in MLLMs
-
Topical Phase Transitions in Artificial Intelligence Research: Large-Scale Evidence and an Early-Warning Signature for Emerging Topics
-
Reasoning for Mobile User Experience with Multimodal LLMs: Task, Benchmark, and Approach
-
Nous: An Attempt to Extract and Inject the Cognition Behind Prediction-Market Behavior
-
Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning
-
LoHoSearch: Benchmarking Long-Horizon Search Agents Beyond the Human Difficulty Ceiling
-
EDEN: A Large-Scale Corpus of Clinical Notes for Italian
-
KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human Difficulty
-
The Order Matters: Sequential Fine-Tuning of LLaMA for Coherent Automated Essay Scoring
-
ParaBridge: Bridging Paralinguistic Perception and Dialogue Behavior in Speech Language Models
-
GaussTrace: Provenance Analysis of 3D Gaussian Splatting Models with Evidence-based LLM Reasoning
-
Density Field State Space Models: 1-Bit Distillation, Efficient Inference, and Knowledge Organization in Mamba-2
-
ConvMemory v2: A Recall-Preserving Top-10 Evidence Reranker for Conversational Memory Retrieval
-
A Navigable Manifold of Hypothesized Consciousness-Spectrum States in Language Model Representations
-
An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs
-
Conservation Laws from Data Symmetry in Neural Networks
-
When Design Rules Break: Benchmark Composition Determines Whether Label Informativeness Predicts GNN Aggregator Choice
-
PL-KKT-hPINN: Enforcing Nonlinear Equality Constraints on Neural Networks via Piecewise-Linear Projection
-
Task Robustness via Re-Labelling Vision-Action Robot Data
-
Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents
-
Rotate2Think: Geometric Priming via Orthogonal Rotation to Improve Language Model Reasoning
-
Less Context, More Accuracy: A Bi-Temporal Memory Engine for LLM Agents Where a Lean Retrieved Context Beats the Full History
-
When Attribution Patching Lies: Diagnosis and a Second-Order Correction
-
Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community
-
FedSteer: Taming Extreme Gradient Staleness in Federated Learning with Corrective Projections and Caching
-
BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression
-
KG-SoftMAP: Soft Knowledge-Graph Priors for Bayesian Network Structure Learning from Sparse Discrete Data
-
Stop Early, Spend Less: Hidden-State Probes as a Practical Recipe for Streaming Moderation of LLM Outputs
-
SD-GRPO: Verifiable Segment Decomposition for Long-Form Vision-Language Generation
-
Conformal Prediction for Neural Operators: Distribution-Free Uncertainty Quantification in Physics Simulation
-
RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning
-
Overcoming Rank Collapse in Feedback Alignment
-
STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios
-
What Fits (Into Few Tokens) Doesn't Overfit: Compression and Generalization in ML Research Agents
-
The Interlocutor Effect: Why LLMs Leak More Personal Data to Agents Than Humans
-
Towards Diverse Scientific Hypothesis Search with Large Language Models
-
Drawing with Strangers: Population Scaling Drives Zero-Shot Mutual Intelligibility in Emergent Sketching
-
Blurry Window Attention
-
Self-EmoQ: Plutchik-Guided Value-based Planning to Drive Streaming Emotional TTS
-
Bellman-Taylor Score Decoding for Markov Decision Processes with State-Dependent Feasible Action Sets
-
The 1st PortraitCraft Challenge: A CVPR 2026 Workshop Competition on Portrait Composition Understanding and Generation
-
Dissect and Prune: Enhancing Robustness in AI-Generated Image Detection
-
POISE: Position-Aware Undetectable Skill Injection on LLM Agents
-
What Makes Video World Model Latents Action-Relevant: Prediction over Reconstruction
-
SRT: Super-Resolution for Time Series via Disentangled Rectified Flow
-
Bridging Traditional Explainability Methods and Multimodal Multilingual Models: An XAI-Based Analysis
-
ZIPP:Zero-shot Image Personalization from Personas
-
FF-JEPA: Long-Horizon Planning in World Models with Latent Planners
-
VESTA: A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents
-
Zero-Shot Learning in Industrial Scenarios: New Large-Scale Benchmark, Challenges and Baseline
-
Traxia: A Framework for Verifiable, Agent-Native Scientific Publishing
-
OmniMem: Perturbation-aware Memory Compression for Streaming Audio-Visual LLMs
-
Observability for Delegated Execution in Agentic AI Systems
-
Code Is More Than Text: Uncertainty Estimation for Code Generation
-
Nonparametric LLM Evaluation from Preference Data
-
Jas: AI-Paired Engineering as a Revival of N-Version Programming
-
TABVERSE: Benchmarking Cross-Format Table Understanding in LLMs and VLMs
-
Implicit Causal Graph Construction in Text via Chain Discovery
-
CoVEBench: Can Video Editing Models Handle Complex Instructions?
-
See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding
-
LEAF: Growing Trees Without Branching for Speech-Aware Large Language Model Post-Training
-
Conan-embedding-v3: Fusing Modality-Specific Models for Omni-Modal Embedding
-
To Nuke or Not to Nuke: LLMs' (Missing) Ethical Reasoning and Actions in a High-Stakes Decision-Making Simulation
-
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
-
Discovering Expert-Level Nash Equilibrium Algorithms with Large Language Models
-
InA-Probe: Instruction-Aware Active Probing for Time Series Forecasting with LLMs
-
DALE-CT: Depth-Aware Foundation Models for Computed Tomography
-
DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning
-
SSAFE: Simple and Strong AI-Generated Image Detection via Frozen Vision Encoders
-
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping
-
HDRAgent: An Agentic Framework for Multi-Exposure HDR Imaging
-
Latent Spatial Memory for Video World Models
-
Training-Free Generalized Few-Shot Segmentation through Open-Vocabulary Semantic Arbitration
-
SoK: Reconstruction Attacks on Synthetic Tabular Data (Insights from Winning the NIST CRC)
-
Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation
-
A large-scale nanocrystal database with aligned synthesis and properties enabling generative inverse design
-
AttentionCap: Transformer Based Capacitance Matrix Learning Toward Full-Chip Extraction
-
The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection
-
BUDDY: BUdget-Driven DYnamic Depth Routing for Adaptive Large Language Model Inference
-
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation
-
Explaining Data Mixing Scaling Laws
-
Stage-1 Controls the Entropy Regime, Not the Outcome
-
NutriMLLM: Multimodal Large Language Models for Dietary Micronutrient Analysis
-
Real-time body pose non-verbal communication with a consistency-based reliability measure
-
SEF-CLGC at SemEval-2026 Task 11: Logical Notation Impact on Language Model Performance
-
NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning
-
GlucoFM-Bench: Benchmarking Time-Series Foundation Models for Blood Glucose Forecasting
-
Accelerating Reproducible Research in Synthetic EHR Generation
-
AI-Driven Test Case Generation from Natural Language Requirements: A Survey of Techniques and Research Gaps
-
Agentic Large Language Models for Automated Structural Analysis of 3D Frame Systems
-
Inside the Visual Mind: Neuroscience-Motivated Concept Circuits for Interpreting and Steering Vision Transformers
-
Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces
-
CULTURESCORE: Evaluating Cultural Faithfulness in Video Generation Models
-
Database Normalization via Dual-LLM Self-Refinement
-
An Expanded Synthetic Conversation Dataset for Multi-Turn Smishing Detection
-
HKVM-RAG: Key-Value-Separated Hypergraph Evidence Organization for Multi-Hop RAG
-
Seeing Without Exposing: Adaptive Privacy Control for Open-World, Context-Hungry MLLMs
-
Evidence Graph Consistency in Retrieval-Augmented Generation: A Model-Dependent Analysis of Hallucination Detection
-
Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns
-
EASE-TTT: Evidence-Aligned Selective Test-Time Training for Long-Context Question Answering
-
LLM Agent-Assisted Reverse Engineering with Quantitative Readability Metrics
-
Evidence-Based Intelligent Diagnostic and Therapeutic Visualization System with Large Language Models: Multi-Turn Interaction and Multimodal Treatment Plan Generation
-
Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software
-
Improving Cross-Lingual Factual Recall via Consistency-Driven Reinforcement Learning
-
M$^3$Exam: Benchmarking Multimodal Memory for Realistic User-Agent Interactions
-
PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration