Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 12 additions & 10 deletions i18n/en.json
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,8 @@
"pageTitle": "NaN Community",
"title": "Join the community while you wait for an inference seat.",
"subtitle": "Discord access, conversations with builders, and a reserved spot in line for GPU access when capacity frees up.",
"priceLabel": "$14.99 / month",
"priceTaxNote": "Taxes included · billed in euros (€14.99) in the EU",
"priceLabel": "14,99€ / month",
"priceTaxNote": "Taxes included · billed in euros from any region",
"includesLabel": "What's included",
"notIncludedLabel": "Not included",
"notIncluded": {
Expand Down Expand Up @@ -302,7 +302,7 @@
"Events, workshops and hackathons",
"Voice in the quarterly model vote"
],
"memberCond": "Month to month, no commitment. Limited spots: when it fills up, waitlist.",
"memberCond": "Month to month, no commitment. Billed in euros from any region. Limited spots: when it fills up, waitlist.",
"memberCta": "Join the waitlist",
"communityIncludes": [
"Discord access",
Expand All @@ -311,13 +311,15 @@
"Events, workshops and hackathons",
"A reserved spot in line for inference"
],
"communityCond": "No commitment. Cancel anytime. Billed in euros (€14.99) in the EU.",
"communityCond": "No commitment. Cancel anytime. Billed in euros from any region.",
"communityCta": "Join the community",
"premiumIncludes": [
"GLM 5.2 — frontier open model with reasoning",
"1,000M token monthly allowance",
"Everything in nan_member · eu",
"Access granted immediately after billing"
"GLM 5.2, frontier open model with reasoning",
"3,000M token allowance per billing period",
"400M tokens per rolling 4h window: the limit a coding agent reaches first",
"500K context · 5 concurrent requests",
"Everything in nan_member",
"Access switches on a few minutes after payment"
],
"premiumCond": "GLM 5.2 premium (200€/mo) is offered to members as seats open. Join the waitlist and we will flag you for this tier.",
"premiumCta": "Join the waitlist",
Expand Down Expand Up @@ -355,7 +357,7 @@
},
{
"q": "Is it really unlimited tokens?",
"a": "On the cluster models, no counter. The frontier ones come with a monthly allowance per member so the plan holds up: 500M tokens on DeepSeek V4-Flash and 1.0B on MiMo v2.5. Those limits are published here, not in fine print. The bill is the fixed subscription price and it doesn’t move."
"a": "On the cluster models, no counter. The frontier ones come with a token allowance per member so the plan holds up: 500M tokens per month on DeepSeek V4-Flash, 1.0B per month on MiMo v2.5, and 3,000M per billing period on GLM 5.2 for members on the premium tier. Those limits are published here, not in fine print. The bill is the fixed subscription price and it doesn’t move."
},
{
"q": "How are the models chosen?",
Expand All @@ -379,7 +381,7 @@
},
{
"q": "What payment methods do you accept?",
"a": "Credit or debit card. Prices include taxes (VAT in the EU, taxes in USA/Latam). No commitment: cancel anytime."
"a": "Credit or debit card. New memberships are billed in euros from any region, taxes included; subscriptions opened earlier in US dollars keep their original price. No commitment: cancel anytime."
},
{
"q": "How do I get in?",
Expand Down
22 changes: 12 additions & 10 deletions i18n/es.json
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,8 @@
"pageTitle": "Comunidad NaN",
"title": "Únete a la comunidad mientras esperas tu plaza de inferencia.",
"subtitle": "Acceso al Discord, charlas con builders y plaza reservada en la cola para acceso a las GPUs cuando se libere capacidad.",
"priceLabel": "$14.99 / mes",
"priceTaxNote": "Impuestos incluidos · en EU se factura en euros (14,99€)",
"priceLabel": "14,99€ / mes",
"priceTaxNote": "Impuestos incluidos · se factura en euros desde cualquier región",
"includesLabel": "Qué incluye",
"notIncludedLabel": "Qué NO incluye",
"notIncluded": {
Expand Down Expand Up @@ -302,7 +302,7 @@
"Eventos, workshops y hackathons",
"Voto en la votación trimestral de modelos"
],
"memberCond": "Mes a mes, sin compromiso. Plazas limitadas: cuando se llena, waitlist.",
"memberCond": "Mes a mes, sin compromiso. Se factura en euros desde cualquier región. Plazas limitadas: cuando se llena, waitlist.",
"memberCta": "Únete a la waitlist",
"communityIncludes": [
"Acceso a Discord",
Expand All @@ -311,13 +311,15 @@
"Eventos, workshops y hackathons",
"Un sitio reservado en la cola de inferencia"
],
"communityCond": "Sin compromiso. Cancela cuando quieras. En EU se factura en euros (14,99€).",
"communityCond": "Sin compromiso. Cancela cuando quieras. Se factura en euros desde cualquier región.",
"communityCta": "Únete a la comunidad",
"premiumIncludes": [
"GLM 5.2 — modelo abierto frontier con razonamiento",
"Asignación mensual de 1.000M de tokens",
"Todo lo incluido en nan_member · eu",
"Acceso activado justo después del cobro"
"GLM 5.2, modelo abierto frontier con razonamiento",
"Asignación de 3.000M de tokens por periodo de facturación",
"400M de tokens por ventana deslizante de 4h: el límite que un agente de código alcanza primero",
"Contexto de 500K · 5 peticiones en paralelo",
"Todo lo incluido en nan_member",
"El acceso se activa pocos minutos después del pago"
],
"premiumCond": "GLM 5.2 premium (200€/mes) se ofrece a miembros según haya plazas. Únete a la waitlist y te marcaremos para este tier.",
"premiumCta": "Únete a la waitlist",
Expand Down Expand Up @@ -355,7 +357,7 @@
},
{
"q": "¿De verdad son tokens ilimitados?",
"a": "En los modelos del cluster, sin contador. Los frontier llevan cuota mensual por miembro para que el plan se sostenga: 500M tokens en DeepSeek V4-Flash y 1.0B en MiMo v2.5. Esos límites están publicados aquí, no en letra pequeña. La factura es el precio fijo de la suscripción y no se mueve."
"a": "En los modelos del cluster, sin contador. Los frontier llevan cuota de tokens por miembro para que el plan se sostenga: 500M tokens al mes en DeepSeek V4-Flash, 1.0B al mes en MiMo v2.5 y 3.000M por periodo de facturación en GLM 5.2 para los miembros del tier premium. Esos límites están publicados aquí, no en letra pequeña. La factura es el precio fijo de la suscripción y no se mueve."
},
{
"q": "¿Cómo se eligen los modelos?",
Expand All @@ -379,7 +381,7 @@
},
{
"q": "¿Qué métodos de pago aceptáis?",
"a": "Tarjeta de crédito o débito. Los precios incluyen impuestos (IVA en la UE, impuestos en USA/Latam). Sin compromiso: cancela cuando quieras."
"a": "Tarjeta de crédito o débito. Las altas nuevas se facturan en euros desde cualquier región, impuestos incluidos; las suscripciones abiertas antes en dólares mantienen su precio original. Sin compromiso: cancela cuando quieras."
},
{
"q": "¿Cómo entro?",
Expand Down
21 changes: 3 additions & 18 deletions src/components/docs/ModelCard.astro
Original file line number Diff line number Diff line change
Expand Up @@ -13,30 +13,15 @@ interface Props {
description: string;
specs: Spec[];
items: string[];
soon?: boolean;
}

const { id, name, tag, leftLabel, rightLabel, description, specs, items, soon } = Astro.props;
const { id, name, tag, leftLabel, rightLabel, description, specs, items } = Astro.props;
const heading = tag ? `${name} - ${tag}` : name;
---

<h2 id={id} data-toc-text={heading}>
{heading}
{
soon && (
<span class="align-middle ml-3 font-mono text-[10px] font-normal uppercase tracking-wider text-amber-300/90 border border-amber-400/30 bg-amber-400/5 rounded-full px-2 py-0.5">
coming soon
</span>
)
}
</h2>
<h2 id={id} data-toc-text={heading}>{heading}</h2>

<div
class:list={[
'grid gap-4 sm:grid-cols-2 mb-14 not-prose',
soon && 'opacity-70',
]}
>
<div class="grid gap-4 sm:grid-cols-2 mb-14 not-prose">
<div class="rounded-xl border border-neutral-800/60 bg-[#0a0a0a] p-6">
<p class="font-mono text-[10px] text-violet-400 uppercase tracking-widest mb-4">
{leftLabel}
Expand Down
48 changes: 46 additions & 2 deletions src/components/docs/RateLimits.astro
Original file line number Diff line number Diff line change
@@ -1,8 +1,14 @@
---
import { env } from 'cloudflare:workers';
import { getRateLimitsConfig } from '../../lib/rateLimits';
import {
formatTokens,
getRateLimitsConfig,
windowedModelBody,
windowedModelHeadline,
} from '../../lib/rateLimits';

const { perKey, tokensPerMinuteByModel, requestsPerMinuteByModel } = getRateLimitsConfig(env);
const { perKey, tokensPerMinuteByModel, requestsPerMinuteByModel, windowedModels } =
getRateLimitsConfig(env);
---

<div class="rounded-xl border border-neutral-800/60 bg-[#0a0a0a] p-6 mb-6">
Expand All @@ -21,6 +27,44 @@ const { perKey, tokensPerMinuteByModel, requestsPerMinuteByModel } = getRateLimi
</dl>
</div>

{
/* The premium models are gated by a sliding window, not by a per-minute
rate. The window is the limit an intensive coding-agent session reaches
first, so it gets its own highlighted block above the per-minute tables
instead of a line of small print underneath them. */
}
{
windowedModels.map((m) => (
<div class="rounded-xl border border-violet-500/30 bg-violet-500/[0.06] p-6 mb-6">
<p class="font-mono text-[10px] text-violet-300 uppercase tracking-widest mb-4">
{m.model} · premium tier limits
</p>
<dl class="grid gap-3 sm:grid-cols-2">
<div class="flex items-baseline justify-between gap-4 font-mono text-xs">
<dt class="text-neutral-500 uppercase tracking-wider">Rolling {m.windowHours}h window</dt>
<dd class="text-neutral-200 text-right">{formatTokens(m.windowTokens)} tokens</dd>
</div>
<div class="flex items-baseline justify-between gap-4 font-mono text-xs">
<dt class="text-neutral-500 uppercase tracking-wider">Allowance / billing period</dt>
<dd class="text-neutral-200 text-right">{formatTokens(m.periodCapTokens)} tokens</dd>
</div>
<div class="flex items-baseline justify-between gap-4 font-mono text-xs">
<dt class="text-neutral-500 uppercase tracking-wider">Context window</dt>
<dd class="text-neutral-200 text-right">{formatTokens(m.contextTokens)} tokens</dd>
</div>
<div class="flex items-baseline justify-between gap-4 font-mono text-xs">
<dt class="text-neutral-500 uppercase tracking-wider">Concurrent requests</dt>
<dd class="text-neutral-200 text-right">{m.maxParallel}</dd>
</div>
</dl>
<p class="mt-4 font-mono text-[11px] leading-relaxed text-neutral-400">
<span class="text-violet-200">{windowedModelHeadline(m)}</span>{' '}
{windowedModelBody(m)}
</p>
</div>
))
}

<div class="rounded-xl border border-neutral-800/60 bg-[#0a0a0a] p-6 mb-6">
<p class="font-mono text-[10px] text-violet-400 uppercase tracking-widest mb-4">
tokens / min por modelo
Expand Down
30 changes: 14 additions & 16 deletions src/components/nan/home/Pricing.astro
Original file line number Diff line number Diff line change
@@ -1,13 +1,23 @@
---
// Datos verificados contra nan.builders: Member EU 70€ · Member USA/Latam $75 · Community $14.99.
// Prices verified against nan.builders: Member 70€ · GLM 5.2 premium 200€ · Community 14,99€.
//
// There is no per-region tier and no per-region currency any more. Every new
// signup is charged in EUR from any region, on BOTH funnels: cloud-api's
// inferencePriceForNewCustomer and communityPriceForNewCustomer each always
// return the EU price. So the `nan_member · usa / latam — $75` tier was removed
// AND community is published in euros here, for the same reason: a prospect
// from outside the EU read a $ amount and landed on a EUR Checkout, a different
// price in a different currency, at the payment step. Legacy USD subscriptions
// still exist and keep their price; that is explained in the payment-methods
// FAQ entry, not here.
import { getLang, useT } from '../../../lib/i18n';

const lang = getLang(Astro.url);
const tt = useT(lang).pricing;
const home = lang === 'es' ? '/es' : '/';

const tiers = [
// GLM 5.2 premium — price swap on the existing nan_member · eu
// GLM 5.2 premium — price swap on the existing nan_member
// subscription (200€/mo, EUR only). Shown first as the flagship tier.
{
name: 'nan_member · glm 5.2 premium',
Expand All @@ -22,7 +32,7 @@ const tiers = [
primary: true,
},
{
name: 'nan_member · eu',
name: 'nan_member',
amount: '70€',
per: tt.perVat,
badge: tt.inferenceIncluded,
Expand All @@ -33,21 +43,9 @@ const tiers = [
href: `${home}#waitlist`,
primary: true,
},
{
name: 'nan_member · usa / latam',
amount: '$75',
per: tt.perTaxes,
badge: tt.inferenceIncluded,
ref: true,
includes: tt.memberIncludes,
cond: tt.memberCond,
cta: tt.memberCta,
href: `${home}#waitlist`,
primary: true,
},
{
name: 'nan_community',
amount: '$14.99',
amount: '14,99€',
per: tt.perTaxes,
badge: null,
ref: false,
Expand Down
18 changes: 12 additions & 6 deletions src/content/docs/api.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -49,7 +49,7 @@ curl https://api.nan.builders/v1/models \

<h2 id="models">GET /v1/models</h2>

Returns the list of available models for your key. Published models: `deepseek-v4-flash`, `mimo-v2.5`, `qwen3.6`, `gemma4`, `qwen3-embedding`, `rerank`, `kokoro`, `whisper`, `flux-2-klein` (includes the `flux-2-klein` image model).
Returns the list of available models for your key. Published models: `deepseek-v4-flash`, `mimo-v2.5`, `qwen3.6`, `gemma4`, `qwen3-embedding`, `rerank`, `kokoro`, `whisper`, `flux-2-klein` (includes the `flux-2-klein` image model). `glm5.2` is served too, and is callable only with a key on the GLM 5.2 premium tier.

### Request

Expand Down Expand Up @@ -86,7 +86,7 @@ curl https://api.nan.builders/v1/models \

<h2 id="chat-completions">POST /v1/chat/completions</h2>

The main chat endpoint. OpenAI Chat Completions compatible. Compatible models: `deepseek-v4-flash`, `mimo-v2.5`, `qwen3.6`, and `gemma4`. `glm5.2` is coming soon and is not callable yet.
The main chat endpoint. OpenAI Chat Completions compatible. Compatible models: `deepseek-v4-flash`, `mimo-v2.5`, `qwen3.6`, `gemma4`, and `glm5.2`. `glm5.2` needs a key on the GLM 5.2 premium tier; the other models are available to every inference member.

<LimitationsCard
title="capabilities by model"
Expand All @@ -107,14 +107,18 @@ The main chat endpoint. OpenAI Chat Completions compatible. Compatible models: `
title: 'gemma4',
body: 'Chat, streaming, vision (image input), reasoning (opt-in).',
},
{
title: 'glm5.2 · premium tier',
body: 'Chat, streaming, tool calling, reasoning trace, 500K token context. Coding and long-horizon agentic tasks. 3,000M token quota per member, reset when your billing period starts, and 400M tokens per rolling 4h window. Requires a key on the GLM 5.2 premium tier.',
},
]}
/>

### Request

| Field | Type | Description |
|---|---|---|
| `model` | string · required | `deepseek-v4-flash`, `mimo-v2.5`, `qwen3.6` or `gemma4`. |
| `model` | string · required | `deepseek-v4-flash`, `mimo-v2.5`, `qwen3.6`, `gemma4` or `glm5.2` (premium tier). |
| `messages` | array · required | List of messages `{ role, content }`. `content` can be a string or an array of parts `[{type:"text",text}, {type:"image_url",image_url:{url}}]` for multimodal input. |
| `max_tokens` | integer · optional | Maximum tokens to generate. |
| `stream` | boolean · optional | Default `false`. If `true`, the response arrives as SSE. |
Expand Down Expand Up @@ -325,14 +329,15 @@ print(data["name"], data["age"])

### Reasoning

All four LLMs generate reasoning and return it in `choices[0].message.reasoning_content`. The control mechanism varies by model:
Every chat LLM generates reasoning and returns it in `choices[0].message.reasoning_content`. The control mechanism varies by model:

| Model | Control |
|---|---|
| `qwen3.6` | `chat_template_kwargs.enable_thinking` · activo por defecto |
| `gemma4` | `chat_template_kwargs.enable_thinking` · desactivado por defecto |
| `deepseek-v4-flash` | `reasoning_effort`: `low` \| `medium` \| `high` · default `medium` |
| `mimo-v2.5` | siempre activo · no configurable por API hoy |
| `glm5.2` | razonamiento activo · sin parámetro publicado para controlarlo |

#### enable_thinking (qwen3.6, gemma4)

Expand Down Expand Up @@ -405,7 +410,7 @@ print(response.choices[0].message.reasoning_content)
print(response.choices[0].message.content)
```

A más `effort`, más tokens dedicados al razonamiento y mejor calidad en problemas complejos — a cambio de latencia y consumo de tu cuota mensual.
A más `effort`, más tokens dedicados al razonamiento y mejor calidad en problemas complejos — a cambio de latencia y consumo de tu cuota de tokens.

#### mimo-v2.5

Expand Down Expand Up @@ -1137,9 +1142,10 @@ Errors follow the standard OpenAI format: HTTP non-2xx status with a JSON body d
|---|---|
| 400 | Parámetro inválido — the body includes `param` with the field that failed (e.g. `prompt`, `n`, `size`, `stream`, `mask` o `image` en los endpoints de imágenes). El filtro de seguridad returns `content_policy_violation`. |
| 401 | `Authorization` header invalid or missing (`invalid_api_key`). |
| 402 | Token allowance for the billing period exhausted on a model that carries one (`monthly_cap_reached`), such as `glm5.2`. Not retryable: the counter goes back to zero when your billing period starts. |
| 403 | Your tier does not have access to the endpoint (`tier_restricted`). Image generation requires inference membership. |
| 404 | Model does not exist (campo `model`, `model_not_found`). |
| 429 | Rate limit exceeded — `rpm_limit` o `max_parallel_requests` (`rate_limit_exceeded`), or monthly quota exhausted (`quota_exceeded` / `insufficient_quota`, como la de 100 requests de imágenes). |
| 429 | Rate limit exceeded — `rpm_limit` o `max_parallel_requests` (`rate_limit_exceeded`), the rolling 4h token budget of `glm5.2` (the body says how much frees up and when), or monthly quota exhausted (`quota_exceeded` / `insufficient_quota`, como la de 100 requests de imágenes). |
| 500 | Internal error (includes upstream model errors). |
| 524 | Timeout (typical with large audios on [/v1/audio/transcriptions](#audio-transcriptions)). |

Expand Down
10 changes: 6 additions & 4 deletions src/content/docs/models.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -64,24 +64,26 @@ with the same `base URL`.
id="glm-5-2"
name="glm5.2"
tag="753B MoE"
soon
leftLabel="text generation & chat — agentic coding"
rightLabel="capabilities"
description="Not available yet: glm5.2 is not published on the community API, so requests naming it return a model error. ~753B parameter MoE model, focused on coding and long-horizon agentic tasks. 256K token context. Tool calling and reasoning (emits a reasoning trace). Text only."
description="Premium tier: callable only with a key on the GLM 5.2 premium membership. ~753B parameter MoE model, focused on coding and long-horizon agentic tasks. 500K token context. Tool calling and reasoning (emits a reasoning trace). Text only. 3,000M token quota per member, and the counter goes back to zero when your billing period starts. Also capped at 400M tokens per rolling 4h window, which is the limit a heavy coding-agent run reaches first."
specs={[
{ label: 'Type', value: 'MoE (~753B total)' },
{ label: 'Quantization', value: 'FP8' },
{ label: 'Attention', value: 'Sparse attention' },
{ label: 'Context', value: '256K tokens' },
{ label: 'Context', value: '500K tokens' },
{ label: 'Input modalities', value: 'text' },
{ label: 'Output modalities', value: 'text' },
{ label: 'Allowance / billing period', value: '3,000M tokens / member' },
{ label: 'Rolling 4h window', value: '400M tokens' },
]}
items={[
'Tool calling (function calling)',
'Reasoning mode (reasoning trace)',
'Coding and long-horizon agentic tasks',
'256K token context',
'500K token context',
'Streaming generation (SSE)',
'Requires the GLM 5.2 premium tier',
]}
/>

Expand Down
Loading
Loading