Designing efficient reward alerts for open-domain query answering is difficult as a result of high-quality responses should concurrently fulfill a number of points of reply high quality which are troublesome to seize with a holistic scalar goal. We introduce a rubric-based reward framework that generates query-specific rubrics grounded in retrieved proof and decomposed into a number of high quality dimensions, offering fine-grained supervision throughout post-training. Averaged throughout three analysis axes (composition, grounding, and instruction-following), our strategy improves over the instruction-tuned baseline by 6.5% and over flat rubric variants by 4%, with constant features throughout all analysis datasets. Conditioning rubrics on retrieved proof improves factual help, whereas decomposing rubrics into quality-specific dimensions additional improves coherence, group, and adherence to question necessities. Our outcomes present that grounded, multi-dimensional rubrics present simpler reward supervision for advanced open-domain query answering.





