Perceive-to-Reason: Fine-Grained Visual Reasoning

This demo uses P2R-8B, a two-stage visual reasoning model based on Qwen3-VL. It first perceives key visual evidence (bounding boxes), then reasons over the highlighted and cropped regions to produce a final answer.

📖 Paper | 🤗 Model | 💻 Code