Entities · Products
Litmaps
32 articles tagged with this entity.
-
Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation
-
Faithful or Findable? Evaluating LLM-Generated Metadata for RDF Dataset Search
-
Integrating knowledge graphs and multilingual scholarly corpora for domain-adaptive LLMs in SSH
-
Mecha-nudges for Machines
-
Rethinking AI-Generated Text Detection: A Strong Baseline and the Distribution-Shift Problem That Remains
-
Attention Limited Reward Learning
-
Human-Centric Reflective Architecture for Human-AI Collaborative Decision-Making
-
Grounding LLM Reasoning under Incomplete Graph Evidence
-
DEEPMED Search: An Open-Source Agentic Platform for Medical Deep Research with Introspective Verification
-
InvestPhilBench: A Multi-Layer Dynamic Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy
-
An Improved Variational Method for Image Denoising
-
MedGuards: Multi-Agent System for Reliable Medical Error Detection and Correction
-
Thinking in Boxes: 3D Editing in Real Images Made Easy
-
Confidence-Aware Automated Assessment of Student-Drawn Scientific Models
-
CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models
-
Simulation of Language Evolution under Regulated Social Media Platforms: A Synergistic Approach of Large Language Models and Genetic Algorithms
-
Dynamic In-Group Persona Generation for Enhancing Human-AI Rapport
-
Evaluative Judgement in Teaching AI-based Translation: A Class-room Case Study of AI-Mediated Translation and Post-Editing
-
Towards Pareto-Optimal Tool-Integrated Agents with Pareto Ranking Policy Optimization
-
Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation
-
SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks
-
Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment
-
Simulating Students' Java Programming Errors with Large Language Models
-
Hasse Diagrams for Attention: A Partial Order Framework for Designing Transformer Masks
-
Towards Personalized Bangla Book Recommendation: A Large-Scale Heterogeneous Book Graph Dataset
-
GIScholarBench: Benchmarking LLM Overconfidence in GIS Research
-
DIYHealth Suite: Dataset, Model, and Benchmark for Health Management at Home
-
Learning to Solve Generative ODEs Beyond the Linear Span
-
ThinkBooster: A Unified Framework for Seamless Test-Time Scaling of LLM Reasoning
-
Evaluating AI-based Scientific Knowledge Synthesis with Epidemiological Systematic Reviews
-
Explicit Evidence Grounding via Structured Inline Citation Generation
-
PaperFlow: Profiling, Recommending, and Adapting Across Daily Paper Streams