HologramHologramModels Hub
Index Sep 17, 2026
Models

TokenRhythm

NeoHorse-1-4B

VerifiedNew4.2B256KTextSafetensors
Address
Identical bytes on
Status
Verified
Trending
#8
Downloads, 30 days
19.8k
Weights
8.4 GB
Sources
2
Revision
Manifest

NeoHorse-1-4B is a 4B causal language model and an initial prototype on the path toward recursive self-improvement (RSI). It is post-trained from Qwen3.5-4B for text-based agent harnesses, tool use, coding, and instruction following.

At a glance

Task
Text generation
Input
text
Output
text
Parameters
4.2B
Architecture
Qwen 3.5 Text
Context
256K tokens
Precision
BF16
Format
Safetensors
Library
transformers
License
Apache 2.0Commercial use
Base model
Fine tuned from Qwen/Qwen3.5-4B
Released
Sep 2026
Updated
Sep 2026
Likes
2,102
Downloads, all time
16,163

Architecture

Layers
32
Hidden size
2,560
Attention
16 heads, grouped query, 4 KV heads
Vocabulary
248,320
Positions
262,144
Tied embeddings
Yes

Family

Models built on NeoHorse-1-4B.

Run it

Loads with Transformers AutoModelForCausalLM and AutoTokenizer, pinned to the indexed revision.

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("TokenRhythm/NeoHorse-1-4B", revision="56f0584bb40578a2c33b1b40a08ccd17243ad710")
processor = AutoTokenizer.from_pretrained("TokenRhythm/NeoHorse-1-4B", revision="56f0584bb40578a2c33b1b40a08ccd17243ad710")

Papers

Spaces

Used in 1 Spaces.

Read the full model card

NeoHorse-1-4B

Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness.

GitHub ModelScope Hugging Face Company Twitter / X License: Apache-2.0

Technical Report

NeoHorse-1-4B is a 4B causal language model and an initial prototype on the path toward recursive self-improvement (RSI). It is post-trained from Qwen3.5-4B for text-based agent harnesses, tool use, coding, and instruction following.

Derived from Qwen/Qwen3.5-4B and fine-tuned by TokenRhythm. This release contains language-model weights only and is repackaged for text-only inference. Vision weights are not included. Repackaging changes configuration and tensor key names, without changing the fine-tuned tensor values.

NeoHorse-1-4B evaluation results

Highlights

  • Path toward RSI: the routing harness assigns tasks to a heterogeneous model pool, records tool interactions and outcomes, estimates capability demand, and uses capability-level feedback to shape the next training mixture. Updated models can return to the harness, closing a prototype evaluation–selection–update loop; extending this loop across successive iterations is the next step toward RSI.
  • Agentic post-training framework: the associated research explores routing-guided curriculum SFT and routing-guided on-policy distillation to turn execution trajectories into training signal while preserving execution and harness context around each response.
  • Data quality: exact and near-duplicate removal, evaluation decontamination, structural validation, six-dimensional semantic evaluation, and subscene-level Scene/Goal/Outcome labeling.
  • Broad gains: 64.87 macro average across ten benchmarks versus 58.94 for Qwen3.5-4B (+5.93).

Model Details

Property Value

Model family NeoHorse Agent-Native Causal Language Model

Parameters Approximately 4B

Base model Qwen3.5-4B

Post-training Routing-guided agentic post-training

Interface Text input and text output

Context length 262,144 natively and extensible up to 1,010,000 tokens.

Weight format / precision Safetensors / BF16

Evaluation

The 4B track compares NeoHorse-1-4B with five representative open-weight models. Results are grouped by capability in the table below. Higher is better; Δ is NeoHorse-1-4B minus Qwen3.5-4B. Bold marks the best available result; underlining marks the second-best.

Benchmark Qwen3.5-4B Gemma-4-E4B-it Nanbeige-4.2-3B Agents-A1-4B Spark-X2.5-4B NeoHorse-1-4B Δ vs Qwen3.5-4B

🤖 Agentic

QwenClawBench

38.47

22.98

40.66

43.16

43.52

44.68

+6.21

WorkBuddy Bench

24.62

11.65

21.03

33.37

26.47

34.41

+9.79

PinchBench

71.19

47.60

66.78

75.07

62.37

77.33

+6.14

VitaBench

21.50

5.00

31.50

39.25

37.00

32.00

+10.50

BFCL v4

61.02

47.18

67.28

46.60

63.71

61.79

+0.77

tau2-Bench

84.29

43.60

85.08

81.00

77.72

88.46

+4.17

💻 Coding

HumanEval

87.20

84.76

98.78

92.68

92.07

96.95

+9.75

LiveCodeBench v6

53.71

52.00

72.50*

56.57

54.86

59.43

+5.72

📚 Instruction Following

IFBench

60.33

40.00

55.00

63.33

73.33

65.33

+5.00

IFEval

87.06

74.68

84.47

83.55

91.13

88.35

+1.29

📊 Overall

Ten-benchmark average

58.94

42.95

62.31

61.46

62.22

64.87

+5.93

* Nanbeige-4.2-3B LiveCodeBench v6 result is reported in the corresponding model's official blog post or technical report.

Reported protocol: SGLang v0.5.17 · temperature=1.0 · top_p=0.95 · top_k=20 · min_p=0.0 · presence_penalty=1.5 · repetition_penalty=1.0 · thinking mode enabled with enable_thinking=true and force_nonempty_content=true. QwenClawBench, WorkBuddy Bench, and tau2-Bench use three runs; PinchBench and VitaBench use one run; the remaining benchmarks follow their official protocols. VitaBench uses the DeepSeek-V4-Flash simulator and judge.

Deployment

The examples below are for self-hosted deployment from a downloaded local checkpoint.

Local checkpoint path

The examples below assume the checkpoint has already been downloaded to local disk. Set MODEL_PATH to the directory containing config.json, tokenizer files, and model weights.

MODEL_PATH="/path/to/NeoHorse-1-4B"

The OpenAI-compatible requests below use the server's --served-model-name (for example, neohorse-1-4b), not the filesystem path.

SGLang

The technical report uses SGLang v0.5.17.

pip install "sglang==0.5.17"
MODEL_PATH="/path/to/NeoHorse-1-4B"
python3 -m sglang.launch_server \
  --model-path "$MODEL_PATH" \
  --served-model-name neohorse-1-4b \
  --host 0.0.0.0 \
  --port 30000 \
  --context-length 262144 \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder

Send an OpenAI-compatible request after the server starts:

curl http://localhost:30000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"neohorse-1-4b","messages":[{"role":"user","content":"Write a Python function that returns the first n Fibonacci numbers."}],"max_tokens":512}'

vLLM

pip install -U vllm
MODEL_PATH="/path/to/NeoHorse-1-4B"
vllm serve "$MODEL_PATH" \
  --served-model-name neohorse-1-4b \
  --host 0.0.0.0 \
  --port 8000 \
  --max-model-len 262144 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder

The server exposes an OpenAI-compatible /v1/chat/completions endpoint. Send a request after the server starts:

curl http://localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"neohorse-1-4b","messages":[{"role":"user","content":"Write a Python function that returns the first n Fibonacci numbers."}],"max_tokens":512}'

The examples use the configured 262,144-token context limit. Actual capacity depends on GPU memory and serving settings; reduce the context limit if needed.

License

NeoHorse-1-4B is released under the Apache License 2.0.

The upstream model is Qwen/Qwen3.5-4B. Its original copyright notice, Copyright 2026 Alibaba Cloud, is retained in the license file. TokenRhythm has modified the model through fine-tuning and repackaging for text-only inference. Modification notices are included in this model card and the released configuration, weight index, and Safetensors metadata.

Citation

@misc{neohorse2026,
  title        = {NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness},
  author       = {NeoHorse Team},
  year         = {2026},
  howpublished = {arXiv preprint},
  eprint       = {2609.08183},
  archivePrefix = {arXiv},
  primaryClass = {cs.CL},
  url          = {https://arxiv.org/abs/2609.08183}
}

For questions or issue reports, use the NeoHorse project repository.

Derived on Sep 16, 2026 from Hugging Face at revision 56f0584b, README.md , config.json .