-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathmodels.yaml
More file actions
394 lines (378 loc) · 15.6 KB
/
Copy pathmodels.yaml
File metadata and controls
394 lines (378 loc) · 15.6 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
schema_version: 3
# These curated entries preserve known profiles, reasoning policies, and
# operator-selected contexts. They take precedence over automatic runtime discovery, but are not
# an allowlist: any other installed Ollama tag can be resolved read-only at run
# time when its effective num_ctx is present in runtime metadata.
models:
laguna-apex-128k:
display_name: Laguna-S-2.1-Uncensored APEX
provider: gx10
runtime: ollama
runtime_model: coder-uncens:latest
runtime_digest: null
quantization: existing-local
context_length: 131072
reasoning_effort: high
reasoning_policy: effort:high
supported_reasoning_policies: ["off", native, effort:high]
role: example-quality-profile
notes:
- A null runtime_digest resolves and records the installed digest at run time.
agent-main:latest:
display_name: Qwen3.6-35B-A3B — 36.0B — Q8_0 — 65536-token context
provider: gx10
runtime: ollama
runtime_model: agent-main:latest
runtime_digest: sha256:0218f872e86baa9c7610509f27db36a7bc52eea7afee24688f81ca74ffcb6c77
source_model: Qwen3.6-35B-A3B
source_version: "3.6"
architecture: qwen35moe
parameter_variant: 36.0B
parameter_count: 35951822704
quantization: Q8_0
context_length: 65536
native_context_length: 262144
reasoning_effort: medium
reasoning_policy: effort:medium
supported_reasoning_policies: ["off", native, effort:medium]
role: curated-runtime-profile
notes:
- Context is operator-configured; the runtime reports a 262144-token native maximum and no stored num_ctx.
coder-max:latest:
display_name: Qwen3-Coder-Next — 79.7B — Q8_0 — 65536-token context
provider: gx10
runtime: ollama
runtime_model: coder-max:latest
runtime_digest: sha256:3f68e12b44eea7c1f464501436bbbe67a8234bf2694efc1d7407ab3df73251b6
source_model: qwen3-coder-next:q8_0
source_version: Next
architecture: qwen3next
parameter_variant: 79.7B
parameter_count: 79674391296
quantization: Q8_0
context_length: 65536
native_context_length: 262144
reasoning_effort: medium
reasoning_policy: effort:medium
supported_reasoning_policies: ["off", native, effort:medium]
role: curated-runtime-profile
notes:
- Context is operator-configured; the runtime reports a 262144-token native maximum and no stored num_ctx.
coder-max-128k:latest:
display_name: Qwen3-Coder-Next — 79.7B — Q8_0 — 131072-token context
provider: gx10
runtime: ollama
runtime_model: coder-max-128k:latest
runtime_digest: sha256:0dbde7ee79ca3ede12b17811e23a4d18e9e61a5dfa261e831cd3f6cdeeaf1803
source_model: qwen3-coder-next:q8_0
source_version: Next
architecture: qwen3next
parameter_variant: 79.7B
parameter_count: 79674391296
quantization: Q8_0
context_length: 131072
native_context_length: 262144
reasoning_effort: medium
reasoning_policy: effort:medium
supported_reasoning_policies: ["off", native, effort:medium]
role: curated-runtime-profile
reviewer-deep:latest:
display_name: GPT-OSS-120B — 116.8B — MXFP4 — 65536-token context
provider: gx10
runtime: ollama
runtime_model: reviewer-deep:latest
runtime_digest: sha256:a951a23b46a1f6093dafee2ea481d634b4e31ac720a8a16f3f91e04f5a40ecd9
source_model: GPT-OSS-120B
source_version: 120B
architecture: gptoss
parameter_variant: 116.8B
parameter_count: 116829156672
quantization: MXFP4
context_length: 65536
native_context_length: 131072
reasoning_effort: medium
reasoning_policy: effort:medium
supported_reasoning_policies: [native, effort:low, effort:medium, effort:high]
role: curated-runtime-profile
notes:
- Context is operator-configured; the runtime reports a 131072-token native maximum and no stored num_ctx.
coder-uncens:latest:
display_name: Laguna-S-2.1-Uncensored APEX — 117.6B — Q6_K — 131072-token context
provider: gx10
runtime: ollama
runtime_model: coder-uncens:latest
runtime_digest: sha256:50197a047af0f7c198929eb80bbd4bb5a44cd43ca878f5241c846ffbf3ee9a7b
source_model: Laguna-S-2.1-Uncensored APEX
source_version: "2.1"
architecture: laguna
parameter_variant: 117.6B
parameter_count: 117561977600
quantization: Q6_K
context_length: 131072
native_context_length: 1048576
reasoning_effort: high
reasoning_policy: effort:high
supported_reasoning_policies: ["off", native, effort:high]
role: curated-runtime-profile
coder-uncens-256k:latest:
display_name: Laguna-S-2.1-Uncensored APEX — 117.6B — Q6_K — 262144-token context
provider: gx10
runtime: ollama
runtime_model: coder-uncens-256k:latest
runtime_digest: sha256:ea34631ee5c00f9acd0182cb9503098e96c551e3841bf537bbc69a4576aea84b
source_model: Laguna-S-2.1-Uncensored APEX
source_version: "2.1"
architecture: laguna
parameter_variant: 117.6B
parameter_count: 117561977600
quantization: Q6_K
context_length: 262144
native_context_length: 1048576
reasoning_effort: high
reasoning_policy: effort:high
supported_reasoning_policies: ["off", native, effort:high]
role: curated-runtime-profile
coder-uncens-qwen:latest:
display_name: Huihui Qwen3-Coder-Next Abliterated — 79.7B — Q8_0 — 131072-token context
provider: gx10
runtime: ollama
runtime_model: coder-uncens-qwen:latest
runtime_digest: sha256:91e5aee4586a37fda420039bfc30293371bab72026a06239e1183339caff2ae0
source_model: huihui_ai/qwen3-coder-next-abliterated:q8_0
source_version: Next
architecture: qwen3next
parameter_variant: 79.7B
parameter_count: 79674391296
quantization: Q8_0
context_length: 131072
native_context_length: 262144
reasoning_effort: medium
reasoning_policy: effort:medium
supported_reasoning_policies: ["off", native, effort:medium]
role: curated-runtime-profile
hermes4-70b:latest:
display_name: Hermes-4-70B — 70.6B — Q4_K_M — 65536-token context
provider: gx10
runtime: ollama
runtime_model: hermes4-70b:latest
runtime_digest: sha256:59c3de79a9d1198dd7065fe1431761274b0bd4fe3582c0c8b59ec044be9627ea
source_model: Hermes-4-70B
source_version: "4"
architecture: llama
parameter_variant: 70.6B
parameter_count: 70553706560
quantization: Q4_K_M
context_length: 65536
native_context_length: 131072
reasoning_effort: medium
reasoning_policy: effort:medium
supported_reasoning_policies: ["off", native, effort:medium]
role: curated-runtime-profile
notes:
- Context is operator-configured; the runtime reports a 131072-token native maximum and no stored num_ctx.
qwen38-q8-medium-262k:
display_name: Qwen3.8 0814 — 27.3B — Q8_0 — 262144-token context
provider: gx10
runtime: ollama
runtime_model: qwen38-q8-262k:latest
runtime_digest: sha256:4ab95509a27d7a3f23dcc612a660858e9f28c1a5322bd9240f34559bcf888988
source_model: qwen3.8:27b-q8_0
source_version: "0814"
architecture: qwen35
parameter_variant: 27.3B
quantization: Q8_0
context_length: 262144
reasoning_effort: medium
reasoning_policy: effort:medium
supported_reasoning_policies: ["off", native, effort:medium]
role: reproduced-standard-profile
default_for_runtime: true
notes:
- The immutable digest reproduces the existing completed standard result.
qwen38-q8-medium-128k:
display_name: Qwen3.8 0814 — 27.3B — Q8_0 — 131072-token context
provider: gx10
runtime: ollama
runtime_model: qwen38-q8-262k:latest
runtime_digest: sha256:4ab95509a27d7a3f23dcc612a660858e9f28c1a5322bd9240f34559bcf888988
source_model: qwen3.8:27b-q8_0
source_version: "0814"
architecture: qwen35
parameter_variant: 27.3B
quantization: Q8_0
context_length: 131072
reasoning_effort: medium
reasoning_policy: effort:medium
supported_reasoning_policies: ["off", native, effort:medium]
role: reproduced-128k-profile
notes:
- This profile shares weights with the 262K entry but freezes a 128K context.
gemma4:31b-it-bf16:
display_name: Gemma 4 31B IT — 31.3B — F16 — 262144-token context
provider: gx10
runtime: ollama
runtime_model: gemma4:31b-it-bf16
runtime_digest: sha256:236d76ae08745dbc143c31b9271b0f25750885199aa6039d0fc0113171606e6d
source_model: gemma4:31b-it-bf16
source_version: "4"
architecture: gemma4
parameter_variant: 31.3B
parameter_count: 31273089132
quantization: F16
context_length: 262144
native_context_length: 262144
template_sha256: sha256:b507b9c2f6ca642bffcd06665ea7c91f235fd32daeefdf875a0f938db05fb315
runtime_capabilities: [completion, thinking, tools, vision]
reasoning_effort: native
reasoning_policy: native
supported_reasoning_policies: ["off", native]
role: curated-runtime-profile
notes:
- Native and off behavior were verified through bounded OpenAI-compatible probes on 2026-08-28; no effort level is inferred.
- The configured smoke on 2026-08-28 preserved an infrastructure diagnostic after both BenchLocal requests exceeded the fixed 180-second case limit; this deployment is currently impractical for GX10 agentic use.
ornith15-q8:latest:
display_name: Ornith-1.5-35B-A3B — 35.5B — Q8_0 — 262144-token context
provider: gx10
runtime: ollama
runtime_model: ornith15-q8:latest
runtime_digest: sha256:c7c57f189918400a4b4193530cceec9136f95c019357de02abc61018f8bad485
source_model: ornith-ai/Ornith-1.5-35B-A3B
source_version: "1.5"
architecture: qwen35moe
parameter_variant: 35.5B
parameter_count: 35505251456
quantization: Q8_0
context_length: 262144
native_context_length: 262144
template_sha256: sha256:f55f52930aa8bf44ab5cb85f99370fcc3c56e9a85640b812086d5330bce5d86b
runtime_capabilities: [completion, thinking, tools, vision]
reasoning_effort: none
reasoning_policy: "off"
supported_reasoning_policies: ["off", native]
role: curated-runtime-profile
notes:
- Native and off behavior were verified through bounded OpenAI-compatible probes on 2026-08-28; no effort level is inferred.
- The intended Hermes deployment disables reasoning; native remains available only as an explicit diagnostic policy.
deepseek-v4-flash:
display_name: DeepSeek V4 Flash — DS4 GGUF + DSpark — 524288-token context
provider: gx10
runtime: ds4
runtime_version: "0.6.5"
runtime_model: deepseek-v4-flash
runtime_digest: sha256:ca22ae2f838e14077c22bc1c1417b71b45b5e5a3687bd96c2ac6e17fdb6261c0
source_model: DeepSeek-V4-Flash
source_version: V4-Flash
architecture: deepseek
parameter_variant: DeepSeek-V4-Flash
quantization: DS4-GGUF
context_length: 524288
reasoning_effort: none
reasoning_policy: "off"
supported_reasoning_policies: ["off"]
endpoint_port: 8000
api_mode: openai-chat-completions
capabilities:
streaming: true
tool_calls: true
deployment_artifacts:
base_gguf_sha256: sha256:ca22ae2f838e14077c22bc1c1417b71b45b5e5a3687bd96c2ac6e17fdb6261c0
dspark_drafter_sha256: sha256:8fa269560dc76fd73e4233ad9b1938b5f65dd363381fd9b1a5c6183f7d12d686
dspark_enabled: true
role: curated-runtime-profile
notes:
- DS4 exposes only /v1/models and /v1/chat/completions to the harness.
- Thinking-disabled control serializes as reasoning_effort=none.
qwen38-flash-next-nvfp4-262k:
display_name: Qwen3.8-Flash-Next — 125B main + 51B PLE / 6B active — NVFP4 — 262144-token context
provider: gx10
runtime: vllm
runtime_version: blazux/qwen3.8-Flash-DGX@209646c
runtime_model: qwen3.8-flash-next
runtime_digest: sha256:56d617b9008714d82b491a73c3fc8f86501a2a4dcd17f0b883359762f1af3e17
runtime_digest_kind: vllm-deployment-provenance
canonical_name: Qwen3.8-Flash-Next
source_model: RadixArk/Qwen3.8-Flash-Next-NVFP4
source_version: Flash-Next
architecture: Qwen3.8-Flash-Next sparse MoE
parameter_variant: 125B main + 51B PLE / 6B active
quantization: NVFP4
context_length: 262144
native_context_length: 262144
reasoning_effort: native
reasoning_policy: native
supported_reasoning_policies: ["off", native]
reasoning_control_profile: qwen-enable-thinking-v1
endpoint_port: 18300
api_mode: openai-chat-completions
capabilities:
streaming: true
tool_calls: true
serving_configuration:
gpu_memory_utilization: 0.75
max_num_seqs: 2
mtp_speculative_tokens: 2
prefix_caching: true
exact_topk: true
prewarm: false
deployment_artifacts:
provenance_contract: vllm-deployment-provenance-v1
checkpoint: RadixArk/Qwen3.8-Flash-Next-NVFP4
checkpoint_revision: 7b719225242aacd3dbd3f9407468c2ee9a9d2594
serving_implementation: blazux/qwen3.8-Flash-DGX
serving_revision: 209646c
serving_revision_full: 209646cd98290035ccbddef29b14c460460a8709
role: curated-runtime-profile
default_for_runtime: false
notes:
- vLLM exposes only /v1/models and /v1/chat/completions to the harness.
- runtime_digest is the SHA-256 of the versioned immutable deployment-provenance descriptor, not a weight-file checksum.
- serving_configuration records the live deployment settings but is intentionally excluded from the immutable provenance digest.
- Native reasoning uses the checkpoint template default; off uses chat_template_kwargs.enable_thinking=false.
- Historical native-thinking deployment with MTP/speculative decoding set to 2; retained for exact attribution of v4 results.
qwen38-flash-next-nvfp4-262k-quality:
display_name: Qwen3.8-Flash-Next — quality deployment (thinking off, MTP off) — NVFP4 — 262144-token context
provider: gx10
runtime: vllm
runtime_version: blazux/qwen3.8-Flash-DGX@209646c
runtime_model: qwen3.8-flash-next
runtime_digest: sha256:56d617b9008714d82b491a73c3fc8f86501a2a4dcd17f0b883359762f1af3e17
runtime_digest_kind: vllm-deployment-provenance
canonical_name: Qwen3.8-Flash-Next
source_model: RadixArk/Qwen3.8-Flash-Next-NVFP4
source_version: Flash-Next
architecture: Qwen3.8-Flash-Next sparse MoE
parameter_variant: 125B main + 51B PLE / 6B active
quantization: NVFP4
context_length: 262144
native_context_length: 262144
reasoning_effort: none
reasoning_policy: "off"
supported_reasoning_policies: ["off", native]
reasoning_control_profile: qwen-enable-thinking-v1
endpoint_port: 18300
api_mode: openai-chat-completions
capabilities:
streaming: true
tool_calls: true
serving_configuration:
gpu_memory_utilization: 0.75
max_num_seqs: 2
mtp_speculative_tokens: 0
prefix_caching: true
exact_topk: true
prewarm: false
kv_cache_dtype: BF16/default
deployment_artifacts:
provenance_contract: vllm-deployment-provenance-v1
checkpoint: RadixArk/Qwen3.8-Flash-Next-NVFP4
checkpoint_revision: 7b719225242aacd3dbd3f9407468c2ee9a9d2594
serving_implementation: blazux/qwen3.8-Flash-DGX
serving_revision: 209646c
serving_revision_full: 209646cd98290035ccbddef29b14c460460a8709
role: curated-runtime-profile
default_for_runtime: true
notes:
- Current quality-oriented deployment identity for gx10-qualification-v5.
- The immutable provenance digest is shared with the historical alias because checkpoint, quantization, runtime model ID, and serving implementation revision are unchanged.
- Mutable deployment provenance separately records thinking off and MTP/speculative decoding off.
- Thinking off serializes as chat_template_kwargs.enable_thinking=false through the vLLM/Qwen reasoning adapter.