- Status
- Verified
- Trending
- #39
- Downloads, 30 days
- 9.9k
- Weights
- 17.9 GB
- Sources
- 2
- Revision
- Manifest
NeoHorse-1-9B is a 9B causal language model and an initial prototype on the path toward recursive self-improvement (RSI). It is post-trained from Qwen3.5-9B for text-based agent harnesses, tool use, coding, and instruction following.
At a glance
- Task
- Text generation
- Input
- text
- Output
- text
- Parameters
- 9B
- Architecture
- Qwen 3.5 Text
- Context
- 256K tokens
- Precision
- BF16
- Format
- Safetensors
- Library
- transformers
- License
- Apache 2.0Commercial use
- Base model
- Fine tuned from Qwen/Qwen3.5-9B
- Released
- Sep 2026
- Updated
- Sep 2026
- Likes
- 717
- Downloads, all time
- 8,736
Architecture
- Layers
- 32
- Hidden size
- 4,096
- Attention
- 16 heads, grouped query, 4 KV heads
- Vocabulary
- 248,320
- Positions
- 262,144
- Tied embeddings
- No
Family
Models built on NeoHorse-1-9B.
Run it
Loads with Transformers AutoModelForCausalLM and AutoTokenizer, pinned to the indexed revision.
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("TokenRhythm/NeoHorse-1-9B", revision="ba5b6e40d88a6ddf4591e176738254a3bc715765")
processor = AutoTokenizer.from_pretrained("TokenRhythm/NeoHorse-1-9B", revision="ba5b6e40d88a6ddf4591e176738254a3bc715765")Papers
- NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing HarnessNeoHorse Team et al., 2026
Spaces
Used in 2 Spaces.
Read the full model card
NeoHorse-1-9B
Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness.
GitHub ModelScope Hugging Face Company Twitter / X License: Apache-2.0
NeoHorse-1-9B is a 9B causal language model and an initial prototype on the path toward recursive self-improvement (RSI). It is post-trained from Qwen3.5-9B for text-based agent harnesses, tool use, coding, and instruction following.
Derived from Qwen/Qwen3.5-9B and fine-tuned by TokenRhythm. This release contains language-model weights only and is repackaged for text-only inference. Vision weights are not included. Repackaging changes configuration and tensor key names, without changing the fine-tuned tensor values.
Highlights
- Path toward RSI: the routing harness assigns tasks to a heterogeneous model pool, records tool interactions and outcomes, estimates capability demand, and uses capability-level feedback to shape the next training mixture. Updated models can return to the harness, closing a prototype evaluation–selection–update loop; extending this loop across successive iterations is the next step toward RSI.
- Agentic post-training framework: the associated research explores routing-guided curriculum SFT and routing-guided on-policy distillation to turn execution trajectories into training signal while preserving execution and harness context around each response.
- Data quality: exact and near-duplicate removal, evaluation decontamination, structural validation, six-dimensional semantic evaluation, and subscene-level Scene/Goal/Outcome labeling.
- Broad gains: 69.04 macro average across ten benchmarks versus 65.60 for Qwen3.5-9B (+3.44).
Model Details
Property Value
Model family NeoHorse Agent-Native Causal Language Model
Parameters Approximately 9B
Base model Qwen3.5-9B
Post-training Routing-guided agentic post-training
Interface Text input and text output
Context length 262,144 natively and extensible up to 1,010,000 tokens.
Weight format / precision Safetensors / BF16
Evaluation
The 9B track compares NeoHorse-1-9B with five representative open-weight baselines: Granite-4.2-8B, Qwen3.5-9B, Ornith-1.5-9B, Gemma-4-12B-it, and Muse-Glimmer-30B. Results cover ten benchmarks and are grouped by capability. Higher is better; Δ is NeoHorse-1-9B minus Qwen3.5-9B. Bold and underline mark the best and second-best results in each benchmark row, respectively; ties share the same formatting.
Benchmark Granite-4.2-8B Qwen3.5-9B Ornith-1.5-9B Gemma-4-12B-it Muse-Glimmer-30B NeoHorse-1-9B Δ vs Qwen3.5-9B
🤖 Agentic
QwenClawBench
37.01
44.04
47.27
43.53
46.11
48.73
+4.69
WorkBuddy Bench
35.07
39.60
29.29
29.65
45.85
40.15
+0.55
PinchBench
56.93
74.55
68.22
58.89
71.35
82.25
+7.70
VitaBench
23.00
31.25
26.75
36.50
48.50
42.25
+11.00
BFCL v4
52.06
64.88
65.03
62.06
53.74
67.43
+2.55
tau2-Bench
62.28
88.04
83.68
59.37
76.64
90.82
+2.78
💻 Coding
HumanEval
96.34
92.68
93.90
100.00
98.17
98.17
+5.49
LiveCodeBench v6
72.00
65.14
47.43
73.14
65.71
65.14
+0.00
📚 Instruction Following
IFBench
78.00
66.33
40.00
77.67
78.67
66.33
+0.00
IFEval
92.98
89.46
71.35
94.27
93.90
89.09
-0.37
📊 Overall
Ten-benchmark average
60.57
65.60
57.29
63.51
67.86
69.04
+3.44
Reported protocol: SGLang v0.5.17 ·
temperature=1.0·top_p=0.95·top_k=20·min_p=0.0·presence_penalty=1.5·repetition_penalty=1.0· thinking mode enabled withenable_thinking=trueandforce_nonempty_content=true. QwenClawBench, WorkBuddy Bench, and tau2-Bench use three runs; PinchBench and VitaBench use one run; the remaining benchmarks follow their official protocols. VitaBench uses the DeepSeek-V4-Flash simulator and judge.
Deployment
The examples below are for self-hosted deployment from a downloaded local checkpoint.
Local checkpoint path
The examples below assume the checkpoint has already been downloaded to local disk. Set MODEL_PATH to the directory containing config.json, tokenizer files, and model weights.
MODEL_PATH="/path/to/NeoHorse-1-9B"
The OpenAI-compatible requests below use the server's --served-model-name (for example, neohorse-1-9b), not the filesystem path.
SGLang
The technical report uses SGLang v0.5.17.
pip install "sglang==0.5.17"
MODEL_PATH="/path/to/NeoHorse-1-9B"
python3 -m sglang.launch_server \
--model-path "$MODEL_PATH" \
--served-model-name neohorse-1-9b \
--host 0.0.0.0 \
--port 30000 \
--context-length 262144 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder
Send an OpenAI-compatible request after the server starts:
curl http://localhost:30000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"neohorse-1-9b","messages":[{"role":"user","content":"Write a Python function that returns the first n Fibonacci numbers."}],"max_tokens":512}'
vLLM
pip install -U vllm
MODEL_PATH="/path/to/NeoHorse-1-9B"
vllm serve "$MODEL_PATH" \
--served-model-name neohorse-1-9b \
--host 0.0.0.0 \
--port 8000 \
--max-model-len 262144 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder
The server exposes an OpenAI-compatible /v1/chat/completions endpoint. Send a request after the server starts:
curl http://localhost:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"neohorse-1-9b","messages":[{"role":"user","content":"Write a Python function that returns the first n Fibonacci numbers."}],"max_tokens":512}'
The example uses the configured 262,144-token context limit. Actual capacity depends on GPU memory and serving settings; reduce the context limit if needed.
License
NeoHorse-1-9B is released under the Apache License 2.0.
The upstream model is Qwen/Qwen3.5-9B. Its original copyright notice, Copyright 2026 Alibaba Cloud, is retained in the license file. TokenRhythm has modified the model through fine-tuning and repackaging for text-only inference. Modification notices are included in this model card and the released configuration, weight index, and Safetensors metadata.
Citation
@misc{neohorse2026,
title = {NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness},
author = {NeoHorse Team},
year = {2026},
howpublished = {arXiv preprint},
eprint = {2609.08183},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2609.08183}
}
For questions or issue reports, use the NeoHorse project repository.
Derived on Sep 16, 2026 from Hugging Face at revision ba5b6e40, README.md , config.json .
