Entities · Products
GotitPub
200 articles tagged with this entity.
-
DiffCVE: Diffusion-based Compressed Video Enhancement
-
Infinite Worlds with Versatile Interactions
-
Named-Entity Recognition in the Crime Domain (CrimeNER): Case Study and Dataset
-
Predicting LLM Safety Before Release by Simulating Deployment
-
AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
-
From Noisy Traces to Root Causes: Structural Trajectory Analysis and Causal Extraction for Agent Optimization
-
Generalist Vision-Language Models for Fast Radio Burst detection: a zero-shot benchmark against a specialized detector
-
Participatory provenance as representational auditing for AI-mediated public consultation
-
Open-Ended Scenario Reasoning for Specialist Model Adaptation
-
A Guiding Framework for K-12 Teachers in Creating AI-powered Learning Technologies through Vibe Coding
-
HumanOmni-Speaker: Identifying Who said What and When
-
i-EXAM: Instructable and Explainable Attack Connectivity Graph Modeler
-
Automated Compliance Mapping in Cloud Security with Domain-Adapted Sentence Transformers
-
Danus: Orchestrating Mathematical Reasoning Agents with Fact-Graph Memory
-
Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic SpeechLLMs
-
Towards Interpretable Foundation Models for Retinal Fundus Images
-
The relationship between reasoning and performance in large language models--o3 (mini) thinks harder, not longer
-
Don't Wait to Reply: Towards Responsive yet Thoughtful Dialogue through Proactive Thinking
-
Reading Between the Dots: Decoding Hidden Computation across Filler Tokens
-
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
-
A Precedent-Guided Co-Scientist for Side-Effect-Aware Drug Redesign
-
SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
-
MAGE: View-guided Point Cloud Completion with Efficient Modality Alignment and Adaptive Geometry Enhancement
-
Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models
-
Builder, Defender, Breaker: The Case Against Removing the Human from the AI-Driven Security Lifecycle
-
Anchored Self-Play for Code Repair
-
TestMate: Test-Time Domain Adaptation Aided by Lightweight Vision Foundation Model
-
When Simpler Is Better: Evaluating Translation Pipelines for Medieval Latin Manuscripts
-
When Do Foundation Models Pay Off? A Break-Even Analysis of Pretrained Time Series Forecasters
-
What is Left for Us? Second Scholarship Against the Degradation of Research by AI
-
DistillH-Mamba: A Hypergraph-Mamba-Based Knowledge Distillation Model for Efficient Impact Fall Detection
-
Trust-Region Noise Search for Black-Box Alignment of Diffusion and Flow Models
-
Weighted Conformal Prediction for Lab-to-Track Thermal Transfer in EV Motorsport Powertrains
-
HAL: Inducing Human-likeness in LLMs with Alignment
-
Transformer Geometry Observatory TGO-II: Representational Similarity Observatory
-
TCG-AR: Real-Time Multi-View Augmented Reality for Trading Card Game Streaming
-
MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering
-
Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning
-
AEGIS: A Multi-Task Joint-Embedding Predictive Architecture for Mammography
-
Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors
-
LV-ROVER: Multi-Stream Tesseract Voting for Maltese Paragraph OCR
-
Hardening x402: PII-Safe Agentic Payments via Pre-Execution Metadata Filtering
-
Planning over MAPF Agent Dependencies via Multi-Dependency PIBT
-
MolSafeEval: A Benchmark for Uncovering Safety Risks in AI-Generated Molecules
-
Readable but Not Controllable: Neuron-Level Evidence for Medical LLM Hallucination
-
Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking
-
Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and Documents
-
ATM: CID-Brokered Pre-Write Admission for Multi-Agent Code Co-Synthesis
-
Multi-Turn Agentic Scientific Literature Search via Workflow Induction
-
A Text-Steerable Instrument for Sketching Procedural Soundscapes via Language Models
-
ZEBRA: Zero-Shot Entropy-Regularized Prompt Learning for Base-to-Novel Generalization in Audio-Language Models
-
Efficient Public Verification of Private ML via Regularization
-
DSIP: A Dynamic Coordination Planner for Signal-Free Intersections using Diffusion-Model-Based Multi-Agent Motion Planning
-
Triospect: A Three-Dimensional Framework for Robust Statistical AI-Generated Text Detection Against Diverse Attacks
-
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
-
Skin-R1: Clinical Knowledge-Guided Dermatological Diagnosis Using Vision-Language Models
-
MixSarc: A Bangla-English Code-Mixed Corpus for Implicit Meaning Identification
-
Manufactured Confidence: How Memory Consolidation Turns Hearsay into Confident Facts
-
Unveiling Novelty Evolution in the field of Library and Information Science in China
-
Evidence-Informed LLM Beliefs for Continual Scientific Discovery
-
Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation
-
PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation
-
Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation
-
GarmentZoom: Generating Zoomable Images from Garment Listings
-
TrafficAlign: Aligning Large Language Models for Traffic Scenario Generation
-
SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models
-
FAIL: Flow Matching Adversarial Imitation Learning for Image Generation
-
Neural Procedural Memory: Empowering LLM Agents with Implicit Activation Steering
-
AnyBody: Free-Form Whole-Body Humanoid Control from Arbitrary Keypoint Guidance
-
UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models
-
GLIP: Graph and LLM Joint Pretraining for Graph-Level Tasks
-
ImprovEvolve: Basin-Hopping Meets LLM-Guided Evolutionary Search
-
SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows
-
Towards Human-Level Book-Writing Capability
-
A Hybrid Framework For Crypto-Ransomware Detection In Enterprise Shared Storage
-
HiComm: Hierarchical Communication for Multi-agent Reinforcement Learning
-
ML-Powered LDAP Reconnaissance Detection using Weak Supervision
-
Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents
-
StableMotion: One-Step Motion Estimation with Diffusion Prior
-
MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations
-
Auto-Configuring Scientific Simulators with Lightweight Coding-Agent Adapters
-
HumanMoveVQA: Can Video MLLMs reason about human movement in videos?
-
LieSolver: PDE-Constrained Learning for IBVPs via Lie Symmetries
-
When Does Personality Composition Matter for Multi-Agent LLM Teams?
-
A Pipeline for Generating Longitudinal Synthetic Clinical Notes Using Large Language Models
-
LithoDreamer: A Physics-Informed World Model for Multi-Stage Computational Lithography
-
AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems
-
PMDformer: Patch-Mean Decoupling Information Transformer for Long-term Forecasting
-
PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing
-
DMuon: Efficient Distributed Muon Training with Near-Adam Overhead
-
Linguistics and Human Brain: A Perspective of Computational Neuroscience
-
How Large Language Models Source Brand Reputation Across Languages and Markets
-
Beyond Visual Forensics: Auditing Multimodal Robustness for Synthetic Medical Image Detection
-
VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks
-
RotRNN: Modelling Long Sequences with Rotations
-
Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages
-
VieSpeaker: A Large-Scale Vietnamese Speaker Recognition Dataset Beyond Visual Dependency
-
REALM: A Unified Red-Teaming Benchmark for Physical-World VLMs
-
Generating adversarial inputs for a graph neural network model of AC power flow
-
Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System
-
SP-Mind: An Autonomous Reasoning Agent for Spatial Proteomics Analysis
-
Data Scale, Not Latency, Shapes Cross-Lingual Encoder Transfer in Streaming ASR
-
Human-like autonomy emerges from self-play and a pinch of human data
-
HEad and neCK TumOR (HECKTOR) 2025: Benchmark of Segmentation, Diagnosis, and Prognosis in Multimodal PET/CT
-
Thinking in Boxes: 3D Editing in Real Images Made Easy
-
PCFootprint: A Large-Scale Dataset and Benchmark for Vectorized Building Footprint Extraction from Aerial LiDAR Point Clouds
-
Calibration Without Comprehension: Diagnosing the Limits of Fine-Tuning LLMs for Vulnerability Detection in Systems Software
-
MIDS: Detecting Stealthy Masquerade and Tampering Attacks on CAN Bus via Bidirectional Mamba
-
As Easy as Rocket Science: Assessing the Ability of Large Language Models to Interpret Negation in Figurative Language
-
GrowthHacker: Automated Off-Policy Evaluation Optimization Using Code-Modifying LLM Agents
-
AI-Driven Assessment of Human Tutors: Linking Training Performance to Real-Life Practice
-
PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation
-
IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages
-
Optimal scenario design for climate emulation
-
Riemannian MeanFlow for One-Step Generation on Manifolds
-
Predictive Analytics in E-Commerce for CustomerBehavior Forecasting using hybrid Ret-DNN withXGBoost Model
-
AI Adoption Across a Multinational Workforce: Sociotechnical Conditions for GenAI Acceptance in Human Resources
-
SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
-
ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents
-
PreAct: Computer-Using Agents that Get Faster on Repeated Tasks
-
Agentic AI-based Framework for Mitigating Premature Diagnostic Handoff and Silent Hallucination in Healthcare Applications
-
NeRD: Neuro-Symbolic Rule Distillation for Efficient Ontology-Grounded Chain-of-Thought in Medical Image Diagnosis
-
Structure-aware Knowledge-guided Heterogeneous Mamba for Zygomaticomaxillary Suture Assessment
-
CPS4: Class Prompt driven Semi-Supervised Spine Segmentation with Class-specific Consistency Constraint
-
MIRAGE: Runtime Scheduling for Multi-Vector Image Retrieval with Hierarchical Decomposition
-
How Post-Training Shapes Biological Reasoning Models
-
MixTeX: Data-Efficient LaTeX OCR via Synthetic Pretraining and Limited Fine-Tuning
-
DifFRACT: Diffusion Feature Reconstruction and Attribution for Circuit Tracing
-
NVMOS: Non-Verbal Vocalization Quality Assessment in Speech
-
Rational Sparse Autoencoder
-
Tight Bounds for Logistic Regression with Large Stepsize Gradient Descent in Low Dimension
-
The Answer Lies Within: Self-Derived Rewards Enable Explainable Relation Extraction
-
Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention
-
TERMS-Bench: Diagnosing LLM Negotiation Agents Beyond Deal Rate
-
Beyond Text-to-SQL: An Agentic LLM System for Governed Enterprise Analytics APIs
-
SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks
-
AIChilles: Automatically Uncovering Hidden Weaknesses in AI-Evolved Systems
-
Simplifying the Modeling of Arbitrary Conditionals in Natural Language
-
Label Shift Aware Adaptation for Online Zero-shot Learning with Contrastive Language-Image Pre-Training (CLIP)
-
A Pragmatic VLA Foundation Model
-
A Robust Point Cloud Analysis Framework Inspired By Primary Visual Cortex
-
Where Black-box Drug-Target Interaction Prediction Models Look: Cross-Method Explainability
-
A Virtuous AI is an Existential Risk
-
MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning Models
-
A Multi-Agent AI System for Automated High School Transcript Processing: Collaborative Document Analysis at Scale
-
Aligned but Stereotypical? How System Prompts Shape Demographic Bias in LLM-Based Text-to-Image Models
-
Listening with Attention: Entropy-Guided Explainability for Transformer-Based Audio Models
-
A New Multi-Domain Benchmark for Micro-Action Recognition and Detection
-
FoleyGenEx: Unified Video-to-Audio Generation with Multi-Modal Control, Temporal Alignment, and Semantic Precision
-
MA-ProofBench: A Two-Tiered Evaluation of LLMs for Theorem Proving in Mathematical Analysis
-
Navigating Gigapixel Pathology Images with Large Multimodal Models
-
VietFashion: Benchmarking Sketch-Text Composed Image Retrieval for Cultural Outfits
-
AfroScope: A Framework for Studying the Linguistic Landscape of Africa
-
M\"OVE: A Holistic LLM Benchmark for the German Public Sector
-
CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges
-
Reconstructing Template-Memorized Images from Natural Prompts
-
WISE: A Long-Horizon Agent in Minecraft with Why-Which Reasoning
-
Brick: Spatial Capability Routing for the Mixture-of-Models (MoM) Paradigm
-
A Quantitative Experimental Repeated Measures Study of Training Dynamics in a Small Llama Style Language Model Under a Compute-Aware Token Budget
-
PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections
-
UniPET: a universal network for high-quality PET image denoising across varied dose reduction factors
-
Swivuriso: The South African Next Voices Multilingual Speech Dataset
-
TENP: Trapezoidal Expert Neuron Pruning For Mixture-of-Experts
-
Spatio-Temporal Attention Graph Neural Network: Explaining Causalities With Attention
-
Expert-Level Crisis Detection in Mental Health Conversations
-
When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models
-
Do VLMs Reason Like Engineers? A Benchmark and a Stage-wise Evaluation
-
Fusing Satellite Imagery and Planimetric Maps for Cross-View Localization
-
++nnU-Net: Scaling nnU-Net with Prefix-Based Data Augmentation
-
MemoVAD: Resource-Efficient Video Anomaly Detection via Dynamic Semantic Memory in Edge Computing Scenarios
-
MatMind: A Structure-Activity Knowledge-Driven Generative Foundation Model for Materials Science
-
GraphLoRA: Structure-Aware Low-Rank Adaptation for Large Language Model Recommendation
-
Momentum for Reasoning: Dense Intrinsic Signals in Policy Optimization
-
Self-Evolving Scientific Agent Discovers Generalizable Physically-Reasoned Fluid Control
-
Contract2Tool: Learning Preconditions and Effects for Reliable Tool-Augmented LLM Agents
-
ArtiFact: A Large-Scale Multi-Modal Cultural Heritage Dataset
-
Distilling Safe LLM Systems via Soft Prompts for On Device Settings
-
Geometry-Aware Uncertainty Quantification via Conformal Prediction on Manifolds
-
Sycophancy as a Multilingual Alignment Failure: How Safety Degrades Across Languages, Topics, and Models
-
EditSR: Enhancing Neural Symbolic Regression via Edit-based Rectification
-
Internalizing Geometric Law: Learning from Solver Residuals for Precision-Critical Generation
-
DIYHealth Suite: Dataset, Model, and Benchmark for Health Management at Home
-
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
-
C$^3$ache: Accelerating World Action Models with Cross Inference Chunk Cache
-
IR-SIM: A Lightweight Skill-Native Simulator for Navigation, Learning, and Benchmarking
-
Order Matters: Unveiling the Hidden Impact of Macro Placement Sequences via Proxy-Guided LLM Evolution
-
Sample-Efficient LLM-Based Detection of Malicious Web Server Logs with Forensically Explainable Reasoning
-
SafeRun: Enabling Determinism in LLM Planning for Running
-
TQA-Bench: Evaluating LLMs for Multi-Table Question Answering
-
Playing Devil's Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy
-
DyCon: Dynamic Reasoning Control via Evolving Difficulty Modeling
-
Training for Technology: Adoption and Productive Use of Generative AI in Legal Analysis
-
Korean Culture into LLM Alignment: Toward Cultural Coherence
-
Quantifying Media Representation Dynamics Across 25 Years of News Reporting on Policing-related Deaths
-
Explicit Evidence Grounding via Structured Inline Citation Generation
-
Phun-Bench: Evaluating LLMs on Phonological Understanding in Chinese
-
Hearing the Unspoken: Language Model Priors for Acoustic Adversarial Attacks
-
Trio: Learning Time-Series Forecasting with Temporal-Spatial-Sample Attention and Structural Causal Priors
-
Re-Centering Humans in LLM Personalization
-
Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation