
- Motivation
- Search-oriented RL improves retrieval but can erode the general reasoning and instruction-following skills expected from universal assistants.
- Method
- Hybrid-OPD couples agentic search RL with cross-domain expert distillation, preserving search specialization while recovering broad capabilities.












