Entities · Products
GSM8K
6 articles tagged with this entity.
-
Reward Granularity in RLVR: Comparing Process and Outcome Reward Structures for Mathematical Reasoning in Small Language Models
-
Training Language Models to Use Prolog as a Tool
-
Weight-Space Geometry of Offline Reasoning Training
-
Pruning via Causal Attribution Preserves Reasoning Performance in Large Language Models
-
Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation
-
Reflection in the Dark: Exposing and Escaping the Black Box in Reflective Prompt Optimization