nyanference Get a key

Inference that lands
in the top corner.

One endpoint, nine models, no seat licences. Chat, embeddings, reranking and transcription — billed by the token and nothing else.

No card required · EU + US regions · OpenAI-compatible schema

bash — nyanference
All systems normal p50 41 ms p99 210 ms 30-day uptime 99.98% tokens today

Endpoints

Everything hangs off one base URL.

If your client already speaks the OpenAI schema, change the base URL and the key. That is the whole migration.

RouteDoesp50Per 1M tokens
POST/v1/chat/completions Streaming and non-streaming chat, tool calls, JSON mode 41 ms$0.60
POST/v1/embeddings 1024-dim vectors, batches up to 2048 inputs 18 ms$0.02
POST/v1/rerank Cross-encoder scoring for retrieval pipelines 26 ms$0.05
POST/v1/audio/transcriptions Speech to text, 40 languages, word timestamps $0.004/min
GET  /v1/models Capabilities, context window and live pricing 9 msfree

Models

Four we host, the rest on request.

Every model runs on our own hardware in Frankfurt and Ashburn. Nothing is proxied to a third party, and nothing you send is trained on.

goat-1

Frontier

Our largest model. Long-context reasoning, tool use, and the only one we let near production code review.

Context256KOutput32KIn / out$0.60 / $2.40

goat-1-mini

Fast

Same tokenizer, a fifth of the cost. Built for classification, extraction and anything you run on every request.

Context128KOutput16KIn / out$0.12 / $0.45

tabby-8b

Open weights

Apache-2.0, fine-tuneable, and downloadable if you would rather run it yourself. We host it because people asked.

Context32KOutput8KIn / out$0.05 / $0.08

Quickstart

Three steps, about four minutes.

1

Request a key

Keys start with nyf_ and are scoped per project. Rotate them from the dashboard whenever you like.

2

Point your client at us

Set the base URL to https://api.nyanference.dev/v1. Existing SDKs work unchanged.

3

Send the first request

Streaming is on by default. Usage shows up in the dashboard within a second or two.


          

The office

Two engineers, four cats, one poster wall.

We are small on purpose. Support is answered by the people who wrote the scheduler, usually within a few hours, occasionally from under a cat.

A tabby cat asleep on a keyboard
Batch scheduler, senior
The poster wall in the office
The wall. It settles most arguments.
A ginger cat sitting on a server rack
Rack thermals, unpaid
A cat watching a monitor
Code review, unsolicited
Framed photograph above a desk
Naming committee, chairman
A cat in a cardboard box beside a desk
Containerisation lead

Pricing

Per token. No minimum, no seats.

Spend caps are hard caps — when you hit one, requests fail with a 429 instead of quietly costing more money.

Free

$0 / month

  • 1M tokens a month
  • 20 requests a minute
  • Community support
  • All models except goat-1

Pay as you go

$0.60 / 1M in

  • Every model
  • 600 requests a minute
  • Email support, same day
  • Spend caps and per-key budgets

Reserved

Talk to us

  • Dedicated capacity
  • Region pinning, EU or US
  • 99.9% written into the contract
  • Invoicing in EUR or USD
Keys are invite-only right now.
We are adding accounts in small batches while the second rack goes in. Mail hello@nyanference.dev with a sentence about what you are building and we will send a key in the next batch — usually inside a week.