Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T05:16:28.008746Z
Paper Citation Record · LEDGER
As of 22 July 2026, this Paper Citation Record lists 31 of 31 outbound references and 16 inbound Pith citation observations for arXiv:2510.18471.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T05:16:28.008746Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-20T06:30:07.809122+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-11T14:30:20.959431Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T12:15:01.137692Z
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cb50c0e3-7c6a-4401-9a99-147dea01ef50 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment OpenCodeReasoning: Advancing Data Distillation for Competitive Coding
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 133023d3-99be-4eae-9140-1fb305078c90 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Evaluating Large Language Models Trained on Code
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 35c381a8-345a-4eea-b64f-bc280f5d75ff · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 9f1ccf0c-4e9f-499f-85d0-45291f2e4ac4 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Process Reinforcement through Implicit Rewards
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 19022fa2-d24f-4201-97be-b9d8355ac07a · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment CodeScore: Evaluating Code Generation by Learning Code Execution
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation f3ba0fca-58f2-4a83-98fe-1e1ceb21b23b · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment A Survey on Code Generation with LLM-based Agents
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 33cd7339-6bfe-4775-ac0e-7e589667d67e · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment ReCode: Reinforcing Code Generation with Reasoning-Process Rewards
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation a92c6490-3bd4-4fd2-9249-c8876bf892ad · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment MiniLLM: On-Policy Distillation of Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 42fdce16-fd98-40e5-8748-dc2aaa83dbe6 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment The False Promise of Imitating Proprietary LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 86d251eb-e19e-42fc-aca2-e7b41d675a63 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Teaching Large Language Models to Reason with Reinforcement Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 4509fc09-955f-49be-9835-cd4b484f275c · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Skywork Open Reasoner 1 Technical Report
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation c92de9ca-b71f-443b-a17b-025fe33c0ba0 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Measuring Coding Challenge Competence With APPS
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation bc0ebe0d-b6de-4509-8776-3e848d148cc1 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Designing and interpreting probes with control tasks
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation feda402c-46a4-464d-9774-51e523e764a1 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Qwen2.5-Coder Technical Report
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 428da1f5-5f6a-42a0-9113-5b1de683fc1d · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation fae20ebe-62cb-41d4-8060-b9a257df0488 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment SEED: customize large language models with sample- efficient adaptation for code generation.CoRR, abs/2403.00046, 2024a
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 883330c4-a3b1-4746-b00f-3c2b18ea02a0 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 498c88f4-de7e-41d7-b0fb-2772d6d3bf35 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment TACO: Topics in Algorithmic COde generation dataset
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 67f68427-9c13-4e8f-8a61-4f441c4caee8 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment GPT-4 Technical Report
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation d75a2d4a-d82e-47e6-924a-1b12a506e19b · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Code Llama: Open Foundation Models for Code
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation f338b61f-d491-4256-adb8-294aac569a5d · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Proximal Policy Optimization Algorithms
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 1ef83e66-5b18-4934-a089-3cab40d05de3 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 7b23b830-ed05-4bf7-99dc-b4222ef13cee · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment HybridFlow: A Flexible and Efficient RLHF Framework
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation bf918166-6fc3-4071-8ec0-d28d694f7339 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment CodeReasoner: Enhancing the Code Reasoning Ability with Reinforcement Learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 7bd90871-6fa5-4941-b7e0-192da2733377 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment CodeBoost: Boosting Code LLMs by Squeezing Knowledge from Code Snippets with RL
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation c8a3f9db-b5ba-4130-b6c0-afbef25450ab · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Co-evolving llm coder and unit tester via reinforcement learning.arXiv preprint arXiv:2506.03136
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation c5509687-8d24-4de0-8b8a-917d97791d98 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation d6eb99f1-bcfd-41d2-bf6a-f13127b68e17 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment A Survey on Knowledge Distillation of Large Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 835bd351-177b-4e26-abc7-1012ea600259 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Here's a step-by-step approach to achieve this:1
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation fb4308cd-a36e-4705-927c-258161822413 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Initialize a Dictionary to Track Blocks: Use a dictionary to map the top-left corner of each 2x2 block to the count of black cells in that block.2
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 815fb496-5bf3-4a86-b4e5-5b659f91ef37 · outbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation d84a2454-7244-45eb-b109-e7f1f8a552cd · inbound
Think Anywhere in Code Generation CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 6a76c5aa-6361-4f47-8d5e-46524980184e · inbound
TestDecision: Sequential Test Suite Generation via Greedy Optimization and Reinforcement Learning CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 830188b8-49f9-4487-9787-53a440b74aeb · inbound
Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation b9f4dcbc-5880-46cb-8691-fba541979f9c · inbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation dfa64b51-5155-4abd-b316-84f35d5c18d6 · inbound
Schedule-and-Calibrate: Utility-Guided Multi-Task Reinforcement Learning for Code LLMs CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation bb4b95a2-b255-4026-908e-42ee84ae03c4 · inbound
Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 7c6a563f-8d8c-4ba8-ad5f-9cff54d5c9dc · inbound
Code as Agent Harness CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation a33f4fe7-b7ed-4d90-9ab7-71c8191ca18e · inbound
Distilling Game Code World Model Generation into Lightweight Large Language Models CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 737d4a5e-111d-4cbb-b10a-0c5d1446197f · inbound
Improving Small Language Models for Code Generation with Reinforcement Learning from Verification Feedback CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 45b56bcc-7699-4426-ac98-e248ca395c1b · inbound
TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation aa580c9a-1676-4f23-8d9d-52405be4a5f2 · inbound
ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 54136184-f7a3-452a-b109-c60542fdb2ef · inbound
Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 1fa12586-08a0-4488-af79-31c964d593d0 · inbound
Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation e2839e89-9d81-4a5b-9f33-b3947090af37 · inbound
From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation a9389d13-50cf-41b5-ad78-a794dfa3a7fa · inbound
When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 401e2898-b4b6-464c-8117-91fcc19a8df3 · inbound
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.