Pistis: Reason deeply, Search broadly, Act with evidence.

Pistis Team, ByteDance

  • Pistis-27B
  • Pistis-9B
  • Qwen3.6-27B
  • Qwen3.5-9B

Radar chart scores

MathVista-mini
Pistis-27B: 87.7; Pistis-9B: 86.8; Qwen3.6-27B: 87.5; Qwen3.5-9B: 85.3.
MM-Vet
Pistis-27B: 77.9; Pistis-9B: 78.1; Qwen3.6-27B: 76.9; Qwen3.5-9B: 75.2.
HallusionBench
Pistis-27B: 69.4; Pistis-9B: 69.3; Qwen3.6-27B: 67.7; Qwen3.5-9B: 66.8.
OCRBench
Pistis-27B: 91.2; Pistis-9B: 91.1; Qwen3.6-27B: 88.0; Qwen3.5-9B: 88.9.
AI2D-test
Pistis-27B: 93.2; Pistis-9B: 92.4; Qwen3.6-27B: 92.6; Qwen3.5-9B: 91.2.
RefCOCO+ testA
Pistis-27B: 94.0; Pistis-9B: 93.9; Qwen3.6-27B: 92.3; Qwen3.5-9B: 89.7.
Charades-STA · 64f
Pistis-27B: 63.2; Pistis-9B: 63.6; Qwen3.6-27B: 56.9; Qwen3.5-9B: 54.1.
MVBench · 8f
Pistis-27B: 70.6; Pistis-9B: 68.9; Qwen3.6-27B: 70.4; Qwen3.5-9B: 67.8.
MMSearch
Pistis-27B: 78.0; Pistis-9B: 72.7; Qwen3.6-27B: 74.7; Qwen3.5-9B: 64.7.
BrowseComp-VL
Pistis-27B: 57.2; Pistis-9B: 53.2; Qwen3.6-27B: 48.4; Qwen3.5-9B: 44.8.
LiveVQA
Pistis-27B: 84.7; Pistis-9B: 82.7; Qwen3.6-27B: 75.0; Qwen3.5-9B: 71.3.
PinchBench
Pistis-27B: 86.5; Pistis-9B: 77.6; Qwen3.6-27B: 85.7; Qwen3.5-9B: 74.8.
TreeBench
Pistis-27B: 59.8; Pistis-9B: 55.3; Qwen3.6-27B: 51.1; Qwen3.5-9B: 52.3.
HRBench4K
Pistis-27B: 91.3; Pistis-9B: 89.0; Qwen3.6-27B: 91.0; Qwen3.5-9B: 87.9.
LogicVista
Pistis-27B: 81.0; Pistis-9B: 70.7; Qwen3.6-27B: 78.3; Qwen3.5-9B: 69.4.
MathVerse-mini
Pistis-27B: 86.9; Pistis-9B: 84.1; Qwen3.6-27B: 85.8; Qwen3.5-9B: 79.8.
Pistis strengthens multimodal reasoning and agentic search across both 27B and 9B models.

Abstract

Pistis is a family of 27B and 9B multimodal models for visual understanding, reasoning, search, and long-horizon interaction. Built on Qwen3.6-27B and Qwen3.5-9B, it provides two complementary variants: Pistis-Thinking and Pistis-Agentic.

Following supervised fine-tuning, our Interleaved Distillation and Reinforcement Learning (IDRL) recipe trains the 9B students by alternating on-policy guidance from the 27B teachers with reward-driven updates. This separation helps sustain useful exploration while improving task performance. Pistis-Auto-Harnessing (PAH) optimizes evidence management, recovery, and execution around a frozen model on development tasks, then freezes the selected harness for inference. Evaluations span multimodal reasoning, agentic search, and Claw-style interaction; matched-budget experiments also show gains from harness optimization without changing model weights.

Method

IDRL: Interleaved Distillation and Reinforcement Learning

Our key idea: learn from the teacher through on-policy distillation, alternate with exploration through reinforcement learning, and aim for stronger generalization and task performance.

Interleaved Distillation and Reinforcement Learning alternates teacher guidance with reward-driven exploration, then finishes with pure RL.

On-policy distillation: the teacher guides the student with top-k knowledge and JSD, helping broaden the policy distribution.

On-policy distillation

The student generates its own responses; the teacher supplies dense, token-level guidance on those same prefixes. Matching their distributions with generalized JSD transfers knowledge while helping preserve useful exploration.

ℓOPD(t)=βDKL(πT∥M)+(1−β)DKL(πθ∥M)\ell_{\mathrm{OPD}}^{(t)}=\beta D_{\mathrm{KL}}(\pi_T\|M)+(1-\beta)D_{\mathrm{KL}}(\pi_\theta\|M)

πT is the teacher, πθ the student, and M their mixture. We use β = 0.5 and the teacher’s top 50 tokens.

Reinforcement learning: rollouts, task rewards, group-relative advantages, and SAPO policy updates sharpen the policy toward better outcomes.

Reinforcement learning

The student samples rollouts, receives task-level rewards, and learns from group-relative advantages through SAPO. Verifiable checks are preferred where available; ineffective tool calls do not receive positive credit merely because the final answer succeeds.

θs+1=θs+η ∇θJRL(θs)\theta_{s+1}=\theta_s+\eta\,\nabla_\theta\mathcal J_{\mathrm{RL}}(\theta_s)

JRL is the SAPO objective; η is the learning rate. This schematic ascent update emphasizes reward-driven improvement.

Why interleave, not sum: alternating OPD and RL steps avoids the instantaneous gradient cross-term in a summed update.

Why interleave, not sum?

Teacher guidance and rewards can point in conflicting directions. IDRL gives them separate updates: OPD re-broadens the policy, while RL reinforces successful behavior. Exactly one objective is active at each step.

Ls={αLOPD,Πs=OPD,−JRL,Πs=RL.\mathcal L_s=\begin{cases}\alpha\mathcal L_{\mathrm{OPD}},&\Pi_s=\mathrm{OPD},\\-\mathcal J_{\mathrm{RL}},&\Pi_s=\mathrm{RL}.\end{cases}

Πs selects the objective; α = 1. Unlike a summed update, each step avoids the immediate cross-term between the two gradients.

PAH: Pistis-Auto-Harnessing

IDRL Model learning

Update the model’s own parameters.

IDRL updates model parametersAlternating OPD and RL updates change the model weights, shown as a changing matrix. This is a conceptual illustration.

PAH Harness evolution

Keep the model fixed. Improve how it acts.

PAH evolves the agent harness around a fixed modelThe central model and its parameters stay fixed. The surrounding code, prompts, evidence management, search skills, and budget control are revised and tested on development tasks. Keep improvements and roll back unsuccessful candidates.
Same model, evolving harness. PAH selects revisions on development tasks, then freezes the harness for inference.
PAH development loop: attribution, proposal, implementation, canary gate, full evaluation, and accept or rollback. The resulting frozen runtime adds a candidate ledger, search skills, and budget-aware checkpoints around Pistis-Agentic.
Figure 4. Development-time harness optimization and the resulting frozen runtime.

Self-evolution through development feedback. An Optimization Agent diagnoses failed trajectories, proposes a targeted code or prompt change, and tests it through a canary gate followed by full development evaluation. It keeps the revision only when the development metric improves; otherwise, it rolls back to the best harness.

h∗=arg max⁡h∈H  Score⁡dev(πfrozen,h;B)h^*=\underset{h\in\mathcal H}{\operatorname{arg\,max}}\;\operatorname{Score}_{\mathrm{dev}}(\pi_{\mathrm{frozen}},h;B)

Select harness h under a fixed model π, tool interface, and interaction budget B. Revisions use 100 development instances. The selected harness is frozen before evaluation on the disjoint 500-sample test set.

A stronger runtime around the same model. The Candidate Ledger preserves evidence and unresolved constraints, Search Skills help recover when search stalls, and checkpoints and budget control reserve room for a final answer. No model weights change, and no Optimization Agent runs at inference time.

Results

Agentic tasks

Pistis-Thinking advances multimodal reasoning; Pistis-Agentic specializes in long-horizon tool use and multimodal search.

Pistis-Thinking-27B

  • Pistis-27B
  • Qwen3.8-27B · xhigh
  • Qwen3.6-27B
MathVista-mini
MME
OCRBench
ChartQA-test
RefCOCO+ testA
Charades-STA · 64f
TACoS · 128f
BLINK

Pistis-Thinking-9B

  • Pistis-9B
  • Qwen3.5-9B
MathVista-mini
MME
OCRBench
ChartQA-test
RefCOCO+ testA
Charades-STA · 64f
TACoS · 128f
BLINK

Pistis-Agentic-27B

  • Pistis-27B
  • Qwen3.8-27B · xhigh
  • Qwen3.6-27B
MMSearch
BrowseComp-VL
VDR-testmini
LiveVQA
PinchBench
TreeBench
MathVerse-mini
LogicVista

Pistis-Agentic-9B

  • Pistis-9B
  • Qwen3.5-9B
MMSearch
BrowseComp-VL
VDR-testmini
LiveVQA
PinchBench
TreeBench
MathVerse-mini
LogicVista

Held-Out Evaluation of the PAH-Optimized Harness

PAH improves how a frozen Pistis-27B-Agentic model uses its tools, then transfers the selected harness without target-specific tuning.

VDR-testmini

500 held-out samples · same 15-interaction limit

Baseline26.8%
PAH-optimized28.6%
+1.8percentage points
Accuracy and interaction budget

Limit = maximum allowed Avg. used = actual interactions.

HarnessLimitAccuracyAvg. used
Same interaction limit
Baseline1526.8%8.68
PAH1528.6%8.60
Other baseline limits · reference only
Baseline1025.2%6.55
Baseline3028.4%11.66

Cross-benchmark transfer

Same frozen harness. No target-specific tuning.

MMSearch
Baseline78.0PAH78.4

+0.4points

BrowseComp-VL
Baseline57.2PAH58.9

+1.7points

LiveVQA
Baseline84.7PAH85.7

+1.0points

The harness is neither revised nor selected on these three benchmarks.

Ablation studies

Starting from the same 9B-Agentic SFT checkpoint, IDRL reaches an overall score of 74.7, compared with 74.1 for sequential and joint training. The improvement is modest but extends to search and Claw-style tasks.

Training strategyChart understandingReal-world perceptionMultimodal reasoningSearch-orientedClaw-styleOverall average
SFT81.077.578.557.177.173.7
Pure RL81.577.879.057.377.574.0
Pure OPD79.678.577.856.274.473.4
Sequential OPD → RL81.278.578.956.577.274.1
Joint RL + OPD81.377.778.957.977.574.1
IDRL82.078.479.458.477.674.7

Citation

If you use Pistis in your research, please cite our technical report.

@misc{chen2026pististechnicalreport,
  title={Pistis Technical Report},
  author={Heyun Chen and Xiaohan Lan and Jiaxi Li and Zhilin Lu and Qi She
          and Weiwen Xu and Fei Yu and Yujie Zhong and Jinghuan Chen
          and Zijian Feng and Siyu Jiao and Yiheng Lin and Xinhao Wang
          and Sihan Yang and Jieyu You and Changbin Zhang and Hengyu Zhang
          and Xudong Zhang and Yunqing Zhao and Shuai Zheng},
  year={2026},
  eprint={2609.28554},
  archivePrefix={arXiv},
  primaryClass={cs.AI},
  url={https://arxiv.org/abs/2609.28554},
}