Occamy-1.0 is a 35B-A3B co-work model continued from Qwen3.6-35B-A3B through full-parameter SFT, uniform model soup, and GRPO/SAO reinforcement learning.
Post-training and evaluation
- A verifier-gated, multi-harness data and training pipeline with immutable provenance and quarantine gates.
- Token-exact proxy/TiTO replay and environment-state reconstruction for long-horizon tool-use trajectories.
- Episode-level credit assignment that remains coherent across context rewrites.
- Frozen, same-harness evaluation for measuring capability and efficiency without silently changing the task contract.
Results
- Raised Combined ClawEval T/C Strict Pass@1/3 from 65.16/73.87% to 77.39/85.93%.
- Reduced tokens per trajectory by 38.6% and trace wall time by 36.1%.
- Reduced timeouts from 10.89% to 4.69% and invalid calls from 3.28% to 1.68%.