flonno.
AN API FOR THE MODELS YOU CHOOSE

One API.
Four models.

Put Kimi, Qwen, Gemma, and NVIDIA models behind one OpenAI-compatible endpoint. Give every application its own key, with only the model access it needs.

Four models One chat endpoint
A REQUEST TO FLONNO01 / 03
POST/v1/chat/completions
curl "https://flonno.com/v1/chat/completions" \
  -H "Authorization: Bearer $FLONNO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3.5-397b-a17b","messages":[{"role":"user","content":"Hello"}]}'
CHOOSE FROM
Moonshot AIAlibaba QwenGoogle GemmaNVIDIA
A SMALLER SURFACE AREA

Make model access
easier to manage.

Keep the model connection familiar while you stay in control of who can use it. Set an access boundary once, then keep your application code focused on the work.

One key per app

Name each key, choose allowed models and endpoints, and set an expiry. Revoke a key without rotating every other app.

01

Keep your client

Send chat completions using a standard OpenAI-compatible request and the public model ID.

02

See request activity

Review model, key, endpoint, status, duration, source address, and token use when the provider reports it.

03
MODEL CATALOG / LIVE PROVIDER PRICES

Find your model.
Know the cost.

Compare input and output prices side by side. Every price is in USD per one million tokens; use the listed model ID in your request.

MODEL / PROVIDERPROVIDER RATE · USD / 1M TOKENS
AVAILABLE MODELINPUTOUTPUT
Moonshot AIKimi K3kimi-k3A capable choice for long-form analysis, planning, and complex questions.$1.0344$5.172
AlibabaQwen 3.5 397B A17Bqwen3.5-397b-a17bA balanced multimodal model for coding, research, and everyday work.$0.15516$1.0344
GoogleGemma 4 31B Turbogemma-4-31b-turboA smaller vision-language model for practical, cost-aware workloads.$0.058764$0.181189
NVIDIANemotron 3 Nano Omni 30Bnemotron-3-nano-omni-30bAn efficient multimodal option for assistants and higher-volume tasks.$0.01098825$0.0438633

These are the model providers’ listed rates, refreshed every 15 minutes from the public catalog. Cached input and other provider billing rules can change the final charge.

Set model access on an API key
FROM KEY TO FIRST REQUEST

Keep the setup
clear and short.

Create a key for an application, pick the access it needs, then call the endpoint with your existing OpenAI client or a direct HTTP request.

Go to API keys
01
Create a separate key

Choose its name, endpoint permissions, models, and optional expiry.

02
Keep the secret server-side

Copy it once into your environment or secret manager.

03
Review and rotate

Check activity, revoke unused access, and replace exposed keys.

base_url = "https://flonno.com/v1"
model = "qwen3.5-397b-a17b"
KNOW WHAT IS RECORDED

Useful logs.
Less content.

Request metadata
Key name, model, endpoint, status, duration, source IP details, and token counts when returned by the provider.

Not written to access logs
Bearer tokens, prompts, and model responses. The full API secret is shown once; only its one-way hash is stored.

Read the full privacy notice
FLONNO / MODEL API

Choose the model.
Keep the access clear.

Create an API key