Entities · Labs
Gotit.pub
287 articles tagged with this entity.
-
From Agentic to Autogenic Network Management for AI-Native 6G and Beyond: A Standards Perspective
-
Ego-Human Motion Prediction with 3D-Aware LLM
-
Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems
-
HIVE: Understanding Post-Hallucination Reasoning in Vision Language Models
-
A Multi-Analyst LLM Pipeline for Auditable Rule Discovery Across 68 Public Physiological Corpora
-
From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents
-
LEMUR 2: Unlocking Neural Network Diversity for AI
-
SPEAR: A Simulator for Photorealistic Embodied AI Research
-
Counterfactual Modeling with Fine-Tuned LLMs for Health Intervention Design and Sensor Data Augmentation
-
PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation
-
DynaKRAG: A Unified Framework for Learnable Evidence Control in Multi-Hop Retrieval-Augmented Generation
-
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents
-
Controlling Tool Use with Heading-Specific Activation Steering
-
Is Domain Adaptation Always Helpful? A Frozen-Backbone Study of Cross-Domain Sentiment Transfer
-
Measuring the practice of shared-decision making (OPTION12): An Investigation into Open-sourced Smaller LLMs (OS-sLLMs) for Better Privacy and Sustainability
-
Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis
-
Base Models Know How to Reason, Thinking Models Learn When
-
Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs
-
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications
-
Scalable Semantic Steering of Embedding Projections
-
ParEVO: Synthesizing Code for Irregular Data: High-Performance Parallelism through Agentic Evolution
-
Convergence Analysis of the ProbAbilistic Gradient Estimator Algorithm for Weakly Convex Finite-Sum Optimization
-
Conditional Diffusion Guided Knowledge Transfer for Multi-Domain Knowledge Graph Completion
-
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning
-
STELLA: Efficient Sensor-to-LLM Translation for On-Device Human Activity Recognition
-
PAST-TIDE: Prototype-Anchored Statement Tuning with Topic-Invariant Normalization for Stance Detection
-
Beyond Static Rules: Automated Discovery of Latent Vulnerabilities in Text-to-SQL
-
GeoSAM-Lite: A Lightweight Foundation Model for Onboard Remote Sensing Segmentation
-
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes
-
Evaluating and Understanding Model Editing for Medical Vision Language Models
-
GLM-5 Serving Parameter Tuning for OpenClaw: Single-Deployment MaaS Inference Optimization for Long-Context Agent Workloads
-
Not Every Sync Is Safe: Calibrated DiLoCo Scheduling for Shared AI Infrastructure
-
From Judgments to Issues: Structured Extraction of Legal Reasoning with Citation-Hallucination Control
-
They Infer What You Meant: Models Represent Communicative Intent More Reliably Than They Act On It
-
Probe, Don't Prompt: A Hidden-State Probe for Metadata Filtering in Multi-Meta-RAG
-
GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech
-
Reduced-Order Models: The Mother of World Models
-
Dictionaries, Not Darwin: Set-Level Selection Beats LLM Evolution in Scientific Equation Discovery
-
Parametric Memory Decoding for Zero-Shot Routing in LoRA-Based External Parametric Memory
-
NKI-Agent: Domain-Specific Fine-Tuning and Agentic Tool Use for Neuron Kernel Generation
-
DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech
-
CONTRA: Red-Teaming Configurations of Personalizable Agents
-
kAgent: An execution-guided crash resolution agent for the Linux kernel
-
Coverage-Controlled Preference Mining from Noisy Claim Verification for Evidence-Grounded Generation
-
When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games
-
Autonomous Information Seeking: A Roadmap for Agentic Recommender Systems
-
WPG-MoE: Weak-Prior-Guided Dense Mixture-of-Experts for User-Level Social Media Depression Detection
-
FormalRx: Rectify and eXamine Semantic Failures in Autoformalization
-
HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control
-
VERITAS: Towards a General-Purpose Replication Tool for Scientific Research
-
Learning Only What Valid Adapters Can Express: Subspace-Constrained Adaptation Against Fine-Tuning Poisoning
-
Silicon Sampling via Cross-Survey Transfer
-
Progressive Disclosure for LLM-Maintained Wiki Knowledge Bases: a Preregistered Ablation
-
GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video
-
Variable Bit-width Quantization: Learning Per-Group Precision for "Bigger-but-Smaller" Language Models
-
Evaluating Agentic Harness Systems for Autonomous Computational Pathology
-
ADVENT: LLM-Driven Automatic Predicate Invention for ILP
-
Beyond Adam: SOAP and Muon for Faster, Label-Efficient Training of Machine Learning Interatomic Potentials
-
On the Limits of Steering Vectors for Preference-Aligned Generation
-
Towards Interactive Global Geolocation Assistant
-
Mapping Text to Multiplex Graph: Prompt Compression as L\'evy Walk-Guided Graph Pruning
-
Class-Grouped Normalized Momentum and Faster Hyperparameter Exploration to Tackle Class Imbalance in Federated Learning
-
Do Newer Lightweight CNNs Perform Better Under Resource Constraints? A Controlled Multigenerational Study of Architecture, Initialization, Training Budget, and Efficiency
-
Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale
-
PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking
-
Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine
-
Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates
-
GEAR-Seg: A Grounded Explainable Agent for Reasoning Segmentation and Data Engine
-
Semantic-Guided Reading Order Reconstruction in Historical Armenian Newspapers with LLMs
-
SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework
-
Real-Time Hard Negative Sampling via LLM-based Clustering for Large-Scale Two-Tower Retrieval
-
Comparative Analysis of Lightweight CNNs for Resource-Constrained Devices: Predictive Performance, Efficiency Trade-offs, and Initialization Effects
-
Towards Robust Driving Perception: A Flexible Scale-Driven Family for Self-Supervised Monocular Depth Estimation
-
Beyond the Prompt: Jailbreaking Function-Calling LLMs via Simulated Moderation Traces
-
Large language models replicate and predict human cooperation across experiments in game theory
-
When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors
-
GR2 Technical Report
-
Indi-RomCoM: Code-Mixed Benchmark for Evaluating LLMs on Romanized Indic-English Instructions
-
HistoriQA-ThirdRepublic: Multi-Hop Question Answering Corpus for Historical Research, Parliamentary Debates from the French Third Republic (1870-1940)
-
Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist
-
Fork-Think with Confidence
-
Bridging Scientific Heritage: An Arabic--Russian Parallel Corpus and LLM Benchmark for Sustainable Knowledge Transfer
-
Can Tabular In-Context Learners Generalize to Biomolecular Property Prediction?
-
An Empirical Study of Security Calibration in Large Language Models for Code
-
Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue
-
Generalizing Numerical Reasoning in Table Data through Operation Sketches and Self-Supervised Learning
-
CLOUDADV: Decision-Aligned Instance Sizing with Zero-Shot Foundation Models under Drift
-
CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
-
Know Before You Fetch: Calibrated Retrieval-Budget Allocation for Retrieval-Augmented Generation
-
Predicting Effects, Missing Distributions: Evaluating LLMs as Human Behavior Simulators in Operations Management
-
Momentum Guidance: Plug-and-Play Guidance for Flow Models
-
Primary ICD Category Prediction using LLM-based Probing
-
Cognitive World Models for Process-Level Social Influence Evaluation
-
Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models
-
StrucTab: A Structured Optimization Framework for Table Parsing
-
Heads, Not Backbones: Output Heads Dominate Architectures on Fat-Tailed Returns
-
DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Model
-
RIPA: Sensory-Vector Prompt Injection Attacks on LLM-Controlled ROS 2 Robots
-
FlatLands: Generative Floormap Completion From a Single Egocentric View
-
CLIMP: Contrastive Language-Image Mamba Pretraining
-
APRIL-MedSeg: A Modular Medical Image Segmentation Toolbox Embracing Modern Paradigms
-
Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation
-
Self-Evolving World Models for LLM Agent Planning
-
Cross-Resolution Semantic Transfer for Robust Text-to-Image Retrieval in Low-Resolution Surveillance
-
FFAvatar: Feed-Forward 4D Head Avatar Reconstruction from Sparse Portrait Images
-
A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis
-
MirrorCode: AI can rebuild entire programs from behavior alone
-
Optimizing Expert-Designed Crystal Graph Networks for Band-Gap Prediction with an Autonomous LLM Research Loop
-
A Machine-Verified Proof of a Quantum-Optimization Conjecture
-
SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics
-
FinInvest-GTCN: Explainable Graph-Temporal-Causal Modeling for Risk-Aware Investment Decision Optimization
-
Fine-Tuning General-Purpose Large Language Models for Agricultural Applications:A Reproducible Framework and Evaluation Protocol Based on Qwen3-8B
-
MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs
-
Latent Bridges for Multi-Table Question Answering
-
HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression
-
CSD: Content-aware Speculative Decoding for Efficient Image Generation
-
Grounded Iterative Language Planning: How Parameterized World Models Reduce Hallucination Propagation in LLM Agents
-
Hybrid Fact-Checking that Integrates Knowledge Graphs, Large Language Models, and Search-Based Retrieval Agents Improves Interpretable Claim Verification
-
VLM-Aware Meta-Optic Front-End Design for Frozen Vision-Language Models
-
DualEval: Joint Model-Item Calibration for Unified LLM Evaluation
-
ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery
-
Rotary Position Encodings for Graphs
-
Investigating LLM's Problem Solving Capability -- a Study on Statics Questions
-
Charting the Growth of Social-Physical HRI (spHRI): A Systematic Review Pipeline Augmented by Small Language Models
-
Robust Onion: Peeling Open Vocab Object Detectors Under Noise
-
SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages
-
ReaORE: Reasoning-Guided Progressive Open Relation Extraction Empowered by Large Reasoning Models
-
The Capability Frontier: Benchmarks Miss 82% of Model Performance
-
Chai: Agentic Discovery of Cryptographic Misuse Vulnerabilities
-
Estimating Uncertainty in Classifier Performance with Applications to Large Language Models and Nested Data
-
RSPC: A Benchmark for Modeling Stress and Psychiatric Conditions in Digitally Mediated Relationships using Psychiatrist Annotations
-
ExTra: Exploratory Trajectory Optimization for Language Model Reinforcement Learning
-
Don't Go Breaking My LLM: The Impact of Pruning Attention Layers on Explanation Faithfulness and Confidence Calibration
-
Do Encoders Suffice? A Systematic Comparison of Encoder and Decoder Safety Judges for LLM Adversarial Evaluation
-
Riazi-8B: An Urdu Large Language Model for Mathematical Reasoning
-
Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models
-
Project Auto-World: Towards Automated Benchmarking of Neural Relational Reasoners
-
Knowledge-Graph Grounding Helps LLMs Only for Out-of-Training Knowledge: A Controlled Study on Clinical Question Answering
-
CANDLE: Character-level Arabic Noise Deduplication using Lightweight Encoder
-
Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning
-
Beyond Logprobs: A Multi-Signal Confidence Engine for LLM-Based Document Field Extraction
-
To Compare, or Not to Compare: On Methodological Practices in Evaluating Social Bias
-
Rule2Text: A Framework for Generating and Evaluating Natural Language Explanations of Knowledge Graph Rules
-
On the Stability of Prompt Ranking in Large Language Model Evaluation
-
BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks
-
RASC+: Retrieval-Constrained LLM Adjudication for Clinical Value Set Authoring
-
L3Cube-MahaPOS: A Marathi Part-of-Speech Tagging Dataset and BERT Models
-
Can Scale Save Us From Plasticity Loss in Large Language Models?
-
LLMs Prompted for Legal Context Object More: Overrefusal from Small On-Premises LLMs in Criminal Legal Context
-
Beyond the Autoregressive Horizon: A Comprehensive Survey of Diffusion Models, World Modelling, and State Space Models for Code
-
Quantifying Prior Dominance in RAG Systems
-
SHERLOC: Structured Diagnostic Localization for Code Repair Agents
-
Contagion Networks: Evaluator Bias Propagation in Multi-Agent LLM Systems
-
AgentArmor: A Framework, Evaluation, \& Mitigation of Coding Agent Failures
-
Low-Burden Data Augmentation for Dysarthric ASR via Zero-Shot Voice Cloning
-
AI Economist Agent: An Agentic Framework for Model-Grounded Economic Analysis with RAG, Knowledge Graphs, and Large Language Models
-
VCG: A Multimodal Retrieval Framework for E-Commerce Video Feeds under Extreme Cold-Start Conditions
-
MiqraBERT: Regression-Based Sentence-BERT Finetuning for Biblical Hebrew Parallel Detection
-
Characterizing Narrative Content in Web-scale LLM Pretraining Data
-
SAM3 Self-Distillation for Fine-Grained GOOSE 2D Semantic Segmentation
-
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
-
Techniques for Peak Memory Reduction for LoRA Fine-tuning of LLMs on Edge Devices
-
Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech
-
Where Does Social Reasoning Come From? Capability Provenance in Language Models
-
Probe-and-Refine Tuning of Repository Guidance for Coding Agents
-
Conflict-Aware Retriever Editing for Knowledge Injection Attacks on LLM-Based RAG Systems
-
Hierarchical Multi-Modal Retrieval for Knowledge-Grounded News Image Captioning
-
Diffusion-Proof: Recipe for Formal Theorem Proving Beyond Auto-Regressive Generation
-
Steerable Cultural Preference Optimization of Reward Models
-
Towards Understanding What State Space Models Learn About Code
-
VISUALSKILL: Multimodal Skills for Computer-Use Agents
-
Dango: A Strictly L1-Only Large Language Model for Studying Second Language Acquisition
-
DreamReasoner-8B: Block-Size Curriculum Learning for Diffusion Reasoning Models
-
Adv-TGD: Adversarial Text-Guided Diffusion for Face Recognition Impersonation Attacks
-
Vibe Coding Ate My Homework: An evaluation of AI approaches to greenfield software engineering and programming
-
SEAGym: An Evaluation Environment for Self-Evolving LLM Agents
-
Agentic Discovery of Non-Canonical Antimicrobial Peptides with AMPGAN v3
-
The Benchmark Illusion: Pruned LLMs Can Pass Multiple Choice but Fail to Answer
-
Reading between the Lines: Leveraging Large Language Models for Global Dementia and Depression Assessment from Clinical Interviews
-
ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling
-
When LLMs Analyze Scars: From Images to Clinically-Meaningful Features
-
Like a Hammer, It Can Build, It Can Break: Large Language Model Uses, Perceptions, and Adoption in Cybersecurity Operations on Reddit
-
Atlas: Orchestrating Heterogeneous Models and Tools for Multi-Domain Complex Reasoning
-
findsylls: A Language-Agnostic Toolkit for Syllable-Level Speech Tokenization and Embedding
-
DICE: Diffusion Large Language Models Excel at Generating CUDA Kernels
-
Graph neural networks at war: integrating cybersecurity and drone intelligence in the Israeli-Iranian conflict
-
Mask Proposal Voting Based on Geodesic Framework for Robust Image Segmentation
-
Deep Residual Injection for Full-Spectrum Forensic Signal Perception in Multimodal Large Language Models
-
Chronological Blindness: Benchmarking Temporal Reasoning in Vision-Language Models with CHRONOSIGHT
-
LLM-Based Visual Explanation Evaluation Framework for Assessing the Explainability of Facial Skin Disease Classification Models
-
Towards Verifiable Agentic Data Science: Solving Irregular TSQA Via Tool-Grounded Reasoning
-
GradPower: Powering Gradients for Faster Language Model Pre-Training
-
Not All Skills Help: Measuring and Repairing Agent Knowledge
-
EIBench: A Simulator-Based Benchmark and Turn-Credit RL for Emotion Management
-
TuneJury: An Open Metric for Improving Music Generation Preference Alignment
-
Risk-Aware LLM Agents for Geospatial Data Retrieval: Design and Preliminary Adversarial Evaluation
-
Gender Differences in AI Literacy Workshop Outcomes and Deepfake Engagement
-
Topological Flow Matching
-
EffGen: Enabling Small Language Models as Capable Autonomous Agents
-
Open-SWE-Traces: Advancing Dual-Mode Multilingual Distillation for Software Engineering Agents
-
OneFocus: Enabling Real-World X-ray Security Screening with a Unified Vision-Language Model
-
Intelligence Is Not the Bottleneck: Validating an LLM First-Pass Manuscript Score Against Peer-Review Outcomes
-
Rapid Poison: Practical Poisoning Attacks Against the Rapid Response Framework
-
Latent Thought Flow: Efficient Latent Reasoning in Large Language Models
-
A Systematic Evaluation of Large Language Models for PTSD Severity Estimation: The Role of Contextual Knowledge and Modeling Strategies
-
FasterPy: An LLM-based Code Execution Efficiency Optimization Framework
-
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
-
PaperJury: Due-Process Review for Bounded LaTeX Revision
-
A Large-Scale Multi-Dimensional Empirical Study of LLMs for Conversation Summarization
-
LLM-as-Code Agentic Programming for Agent Harness
-
SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity
-
Human genetic evidence is associated with drug approval across therapeutic areas: an observational analysis of 26,278 target-disease pairs with temporal validation and feature ablation
-
Let LLMs Judge Each Other: Multi-Agent Peer-Reviewed Reasoning for Medical Question Answering
-
Koshur Diacritizer: A Byte-Level Sequence-to-Sequence Model for Kashmiri Diacritic Restoration
-
PVminerLLM2: Improving Structured Extraction of Patient Voice via Preference Optimization
-
Recipe-Controlled Decoder Audit for Structural Knowledge-Graph Completion
-
From Prompts to Responses: Dual-Sided Data Leakage and Defense in Split Large Language Models
-
MeEvo: Metacognitive Evolution Combined with Natural Evolution for Automatic Heuristic Design
-
Clay-CNN Hybrids: Leveraging Geo-Foundational Models as Auxiliary Context for Landslide Detection
-
DLawBench: Evaluating LLMs Through Multi-Turn Legal Consultation
-
Harsher on Male? Evaluating LLMs on Gender-Asymmetric Moral Framing Across Diverse Conflict Scenarios
-
GRIP: Feedback-Guided Prompt Retrieval for Large Multimodal Models
-
Shopping Reasoning Bench: An Expert-Authored Benchmark for Multi-Turn Conversational Shopping Assistants
-
FENCE: A Financial and Multimodal Jailbreak Detection Dataset
-
One Polluted Page Is Enough: Evaluating Web Content Pollution in Generative Recommenders
-
IVIE: A Neuro-symbolic Approach to Incremental and Validated Generation of Interactive Fiction Worlds
-
HyperTool: Beyond Step-Wise Tool Calls for Tool-Augmented Agents
-
DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasks
-
Agents-K1: Towards Agent-native Knowledge Orchestration
-
LLM-Powered Personalized Glycemic Assessment in Type 2 Diabetes with Wearable Sensor Data
-
Stubborn: A Streamlined and Unified Reinforcement Learning Framework for Robust Motion Tracking and Fall Recovery for Humanoids
-
Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It
-
N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization
-
Calibrating Overconfidence Without Sacrificing Confidence: Probe-Conditioned Head Intervention for LLMs
-
Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation
-
PhantomBench: Benchmarking the Non-existential Threat of Language Models
-
TinyTroupe: An LLM-powered Multiagent Persona Simulation Toolkit
-
IDP-Bench: Benchmarking ability of LLMs to protect personal information in interdependent privacy contexts
-
When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff
-
3SPO: State-Score-Supervised Policy Optimization for LLM Agents
-
Generative Explainability for Next-Generation Networks: LLM-Augmented XAI with Mutual Feature Interactions
-
Speech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and Translation
-
UPLOTS: A Unified Pretrained Language Model for Constrained Time-series Generation
-
Predicting Future Behaviors in Reasoning Models Enables Better Steering
-
EstRTL: Functional Estimation Guided RTL Code Generation
-
Mix, Don't Pick: Why Synthetic Corpus Composition Matters for Time Series Foundation Model Pretraining
-
FADA: Accessible fetal ultrasound interpretation and annotation with a selectively distilled unified vision-language model
-
Liberating LLM Capabilities in Full-Duplex Speech Models
-
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research
-
Principled Agent Debate: Adversarial Arbitration for Sycophancy Reduction in Large Language Models
-
Post-training is (Massive) Supervised Learning
-
MASS: Deep Research for Social Sciences with Memory-Augmented Social Simulation
-
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
-
Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs
-
DN-Hypo-Pipeline: An AI-Driven Workflow for Hypothesis Generation via Large Language Models and Scientific Explanations
-
Why Limit the Residual Stream to Layers and Not Tokens? Persistent Memory for Continuous Latent Reasoning
-
POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
-
Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks
-
Efficient Skill Grounding via Code Refactoring with Small Language Models
-
GIFT: LLM-Guided State-Reward Interface for Financial Reinforcement Learning
-
OmniTryOn: Video Try-On Anything at Once!
-
Claude Code-Driving Scenario Mining for the Argoverse 2 Challenge
-
POTATR: A Lightweight Image-to-Graph Model for Page-Level Table Extraction
-
Visual Template Inference for Data Extraction from Documents
-
Sparse Memory Finetuning as a Low-Forgetting Alternative to LoRA and Full Finetuning
-
FADTI: Fourier and Attention Driven Diffusion for Multivariate Time Series Imputation
-
APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music
-
VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation
-
Contemporary AI lacks the imagination to diverge or negate in science
-
FMplex: Model Virtualization for Serving Extensible Foundation Models
-
ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning
-
MalTree: Tracing Malware Evolution from Embeddings at Scale
-
ShallowBench: Benchmarking Generative Drug Design Models on Shallow-Pocket Targets
-
OpenHalDet: A Unified Benchmark for Hallucination Detection across Diverse Generation Scenarios
-
SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating
-
GP-Adapter: Gaussian Process CLIP-Adapter for Few-Shot Out-of-Distribution Detection
-
When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations
-
MADE: Beyond Scoring via a Multilingual Agentic Diagnosing Engine for Fine-Grained Evaluation Insights
-
Beyond Rubrics: Exploration-Guided Evaluation Skills for Reward Modeling
-
OpenGlass: Open-Source Smart Glasses for On-Device Event-Based Gesture Recognition
-
MacArena: Benchmarking Computer Use Agents on an Online macOS Environment
-
MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval
-
GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning
-
Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle
-
ActionMap: Robot Policy Learning via Voxel Action Heatmap
-
Are Large Language Models Suitable for Graph Computation? Progress and Prospects
-
Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents