Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals Paper • 2607.11505 • Published 10 days ago • 18
Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly? Paper • 2606.26428 • Published 29 days ago • 17
AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systems Paper • 2605.27466 • Published May 26 • 8
kairawal/Llama-3.2-1B-Instruct-DA-SynthDolly-r16alpha32-E8-S73 Text Generation • 1B • Updated May 18 • 4 • 1