- Pistis-27B
- Pistis-9B
- Qwen3.6-27B
- Qwen3.5-9B
Radar chart scores
- MathVista-mini
- Pistis-27B: 87.7; Pistis-9B: 86.8; Qwen3.6-27B: 87.5; Qwen3.5-9B: 85.3.
- MM-Vet
- Pistis-27B: 77.9; Pistis-9B: 78.1; Qwen3.6-27B: 76.9; Qwen3.5-9B: 75.2.
- HallusionBench
- Pistis-27B: 69.4; Pistis-9B: 69.3; Qwen3.6-27B: 67.7; Qwen3.5-9B: 66.8.
- OCRBench
- Pistis-27B: 91.2; Pistis-9B: 91.1; Qwen3.6-27B: 88.0; Qwen3.5-9B: 88.9.
- AI2D-test
- Pistis-27B: 93.2; Pistis-9B: 92.4; Qwen3.6-27B: 92.6; Qwen3.5-9B: 91.2.
- RefCOCO+ testA
- Pistis-27B: 94.0; Pistis-9B: 93.9; Qwen3.6-27B: 92.3; Qwen3.5-9B: 89.7.
- Charades-STA · 64f
- Pistis-27B: 63.2; Pistis-9B: 63.6; Qwen3.6-27B: 56.9; Qwen3.5-9B: 54.1.
- MVBench · 8f
- Pistis-27B: 70.6; Pistis-9B: 68.9; Qwen3.6-27B: 70.4; Qwen3.5-9B: 67.8.
- MMSearch
- Pistis-27B: 78.0; Pistis-9B: 72.7; Qwen3.6-27B: 74.7; Qwen3.5-9B: 64.7.
- BrowseComp-VL
- Pistis-27B: 57.2; Pistis-9B: 53.2; Qwen3.6-27B: 48.4; Qwen3.5-9B: 44.8.
- LiveVQA
- Pistis-27B: 84.7; Pistis-9B: 82.7; Qwen3.6-27B: 75.0; Qwen3.5-9B: 71.3.
- PinchBench
- Pistis-27B: 86.5; Pistis-9B: 77.6; Qwen3.6-27B: 85.7; Qwen3.5-9B: 74.8.
- TreeBench
- Pistis-27B: 59.8; Pistis-9B: 55.3; Qwen3.6-27B: 51.1; Qwen3.5-9B: 52.3.
- HRBench4K
- Pistis-27B: 91.3; Pistis-9B: 89.0; Qwen3.6-27B: 91.0; Qwen3.5-9B: 87.9.
- LogicVista
- Pistis-27B: 81.0; Pistis-9B: 70.7; Qwen3.6-27B: 78.3; Qwen3.5-9B: 69.4.
- MathVerse-mini
- Pistis-27B: 86.9; Pistis-9B: 84.1; Qwen3.6-27B: 85.8; Qwen3.5-9B: 79.8.
Abstract
Pistis is a family of 27B and 9B multimodal models for visual understanding, reasoning, search, and long-horizon interaction. Built on Qwen3.6-27B and Qwen3.5-9B, it provides two complementary variants: Pistis-Thinking and Pistis-Agentic.
Following supervised fine-tuning, our Interleaved Distillation and Reinforcement Learning (IDRL) recipe trains the 9B students by alternating on-policy guidance from the 27B teachers with reward-driven updates. This separation helps sustain useful exploration while improving task performance. Pistis-Auto-Harnessing (PAH) optimizes evidence management, recovery, and execution around a frozen model on development tasks, then freezes the selected harness for inference. Evaluations span multimodal reasoning, agentic search, and Claw-style interaction; matched-budget experiments also show gains from harness optimization without changing model weights.
Method
IDRL: Interleaved Distillation and Reinforcement Learning
Interleaved Distillation and Reinforcement Learning alternates teacher guidance with reward-driven exploration, then finishes with pure RL.
On-policy distillation
The student generates its own responses; the teacher supplies dense, token-level guidance on those same prefixes. Matching their distributions with generalized JSD transfers knowledge while helping preserve useful exploration.
πT is the teacher, πθ the student, and M their mixture. We use β = 0.5 and the teacher’s top 50 tokens.
Reinforcement learning
The student samples rollouts, receives task-level rewards, and learns from group-relative advantages through SAPO. Verifiable checks are preferred where available; ineffective tool calls do not receive positive credit merely because the final answer succeeds.
JRL is the SAPO objective; η is the learning rate. This schematic ascent update emphasizes reward-driven improvement.
Why interleave, not sum?
Teacher guidance and rewards can point in conflicting directions. IDRL gives them separate updates: OPD re-broadens the policy, while RL reinforces successful behavior. Exactly one objective is active at each step.
Πs selects the objective; α = 1. Unlike a summed update, each step avoids the immediate cross-term between the two gradients.
PAH: Pistis-Auto-Harnessing
IDRL Model learning
Update the model’s own parameters.
PAH Harness evolution
Keep the model fixed. Improve how it acts.

Self-evolution through development feedback. An Optimization Agent diagnoses failed trajectories, proposes a targeted code or prompt change, and tests it through a canary gate followed by full development evaluation. It keeps the revision only when the development metric improves; otherwise, it rolls back to the best harness.
Select harness h under a fixed model π, tool interface, and interaction budget B. Revisions use 100 development instances. The selected harness is frozen before evaluation on the disjoint 500-sample test set.
A stronger runtime around the same model. The Candidate Ledger preserves evidence and unresolved constraints, Search Skills help recover when search stalls, and checkpoints and budget control reserve room for a final answer. No model weights change, and no Optimization Agent runs at inference time.
Results
Agentic tasks
Pistis-Thinking advances multimodal reasoning; Pistis-Agentic specializes in long-horizon tool use and multimodal search.
Pistis-Thinking-27B
- Pistis-27B
- Qwen3.8-27B · xhigh
- Qwen3.6-27B
Pistis-Thinking-9B
- Pistis-9B
- Qwen3.5-9B
Pistis-Agentic-27B
- Pistis-27B
- Qwen3.8-27B · xhigh
- Qwen3.6-27B
Pistis-Agentic-9B
- Pistis-9B
- Qwen3.5-9B
Held-Out Evaluation of the PAH-Optimized Harness
PAH improves how a frozen Pistis-27B-Agentic model uses its tools, then transfers the selected harness without target-specific tuning.
VDR-testmini
500 held-out samples · same 15-interaction limit
Limit = maximum allowed Avg. used = actual interactions.
| Harness | Limit | Accuracy | Avg. used |
|---|---|---|---|
| Same interaction limit | |||
| Baseline | 15 | 26.8% | 8.68 |
| PAH | 15 | 28.6% | 8.60 |
| Other baseline limits · reference only | |||
| Baseline | 10 | 25.2% | 6.55 |
| Baseline | 30 | 28.4% | 11.66 |
Cross-benchmark transfer
Same frozen harness. No target-specific tuning.
MMSearch
+0.4points
BrowseComp-VL
+1.7points
LiveVQA
+1.0points
The harness is neither revised nor selected on these three benchmarks.
Ablation studies
Starting from the same 9B-Agentic SFT checkpoint, IDRL reaches an overall score of 74.7, compared with 74.1 for sequential and joint training. The improvement is modest but extends to search and Claw-style tasks.
| Training strategy | Chart understanding | Real-world perception | Multimodal reasoning | Search-oriented | Claw-style | Overall average |
|---|---|---|---|---|---|---|
| SFT | 81.0 | 77.5 | 78.5 | 57.1 | 77.1 | 73.7 |
| Pure RL | 81.5 | 77.8 | 79.0 | 57.3 | 77.5 | 74.0 |
| Pure OPD | 79.6 | 78.5 | 77.8 | 56.2 | 74.4 | 73.4 |
| Sequential OPD → RL | 81.2 | 78.5 | 78.9 | 56.5 | 77.2 | 74.1 |
| Joint RL + OPD | 81.3 | 77.7 | 78.9 | 57.9 | 77.5 | 74.1 |
| IDRL | 82.0 | 78.4 | 79.4 | 58.4 | 77.6 | 74.7 |
Citation
If you use Pistis in your research, please cite our technical report.
@misc{chen2026pististechnicalreport,
title={Pistis Technical Report},
author={Heyun Chen and Xiaohan Lan and Jiaxi Li and Zhilin Lu and Qi She
and Weiwen Xu and Fei Yu and Yujie Zhong and Jinghuan Chen
and Zijian Feng and Siyu Jiao and Yiheng Lin and Xinhao Wang
and Sihan Yang and Jieyu You and Changbin Zhang and Hengyu Zhang
and Xudong Zhang and Yunqing Zhao and Shuai Zheng},
year={2026},
eprint={2609.28554},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2609.28554},
}
Pistis:
