Nepal’s first LLM router

M82 Router: One API key.Every model.Billed in Rs.

An OpenAI-compatible AI router for developers in Nepal. Access multiple models and see the exact rupee cost of every request.

50,000 tokens in, 5,000 out on Claude Haiku 5.5 by Anthropic: Rs. 1.28All 41 models

M82 Router · AI infrastructure

An LLM router built for Nepal.

M82 Router connects your application to multiple large language models through one OpenAI-compatible API. Choose a model from the catalog, send your request with an M82 API key, and see its cost in Nepali rupees.

What is an LLM router?

An LLM router, also called an AI router, gives your application a common API for different large language models. With M82 Router, you select the model for each request and can switch between supported models without managing a separate provider key for each one.

Read the API integration guide for curl, Python, Node.js and OpenCode setup.

LLM API access with local billing

For developers building from Nepal, M82 Router brings model prices and request costs into NPR, with eSewa as the top-up method. Account access and payments depend on the current beta availability shown on this site.

Browse models and rupee prices or compare published model usage to choose a model for your application.

How it works

From eSewa to your first response in three steps.

  1. 01Wallet

    Top up with eSewa

    Add credit in Nepali rupees. No international card, no dollar account.

    Example amount.

  2. 02Key

    Create an API key

    Your key is shown once. We keep only a fingerprint of it, never the key itself.

    Masked example, not a real key.

  3. 03Client

    Change your base URL

    Keep your OpenAI SDK. Swap the base URL and the key, and nothing else.

    Works with OpenAI's Python and Node SDKs, curl and OpenCode.

Billing

Priced before it runs. Settled exactly once.

Every request holds credit from your balance before it runs. When the reply ends, you pay once for what it actually used.

  1. 01

    Reserve

    Hold an estimate of the cost: what you send, plus your max_tokens.

  2. 02

    Run

    The model replies, streamed or all at once. Nothing touches your wallet.

  3. 03

    Settle

    The real usage is priced once: one charge, and any unused hold is released.

One request on gpt-oss-20breq_01M4M8TAVWH17EPD370FW04AGJ
Example
Ledger entries for one example request, with the available balance after each step
StepEntryAmount, Rs.Available, Rs.
Startwallet balance—500.000000000000
01Reservereservation−0.067645857888499.932354142112
02Runusage1,200 in · 350 outno change
03Settleusage_charge−0.009037829745499.990962170255
reservation_release+0.058608028143
Charged onceRs. 0.009037829745

402If the hold doesn't fit your available balance, the request stops here and is never sent to a model.

Example: a 4,800-byte request on gpt-oss-20b with max_tokens 4,096, starting from a Rs. 500 balance. Settle writes both of its lines together.

Why M82 Router

No dollar card. No surprise bills.

The exact cost on every response

Normal responses carry m82.cost_npr and a request ID that matches one line in your Usage page and ledger.

"usage": { "prompt_tokens": 1200, "completion_tokens": 350 },
"m82": {
  "request_id": "req_01M4M8TAVWH17EPD370FW04AGJ",
  "cost_npr": "0.009037829745"
}

Example: a short reply on gpt-oss-20b.

Pay with eSewa

Credit in Nepali rupees. No dollar card, so no card loading charge and no exchange spread.

Pay before you use

Prepaid. A request your balance can't cover gets a 402 and never reaches the model.

  • model, tokens, cost, request ID
  • prompts
  • completions

No prompt storage

We don't store or log what you send or what comes back. Only the metadata you need to reconcile a charge.

Fair limits, per key and account

Each key gets about 10,00,000 tokens a minute and 8 requests at once; your account 300 requests a minute in total. Over the limit, a 429 with Retry-After.

Integration

Change two lines. Keep your client.

M82 Router speaks the OpenAI Chat Completions format, including streaming and tool calls. We test it end to end with curl and OpenCode.

import osfrom openai import OpenAI client = OpenAI(removed:     base_url="https://api.openai.com/v1",removed:     api_key=os.environ["OPENAI_API_KEY"],added:     base_url="https://api.m82router.com/v1",added:     api_key=os.environ["M82_API_KEY"],    max_retries=0,) reply = client.chat.completions.create(    model="openai/gpt-oss-20b",    messages=[{"role": "user", "content": "Say hello in Nepali."}],)print(reply.choices[0].message.content)

Keep your key in an environment variable or a secret manager, never in browser code.

Response200 OK
{
  "object": "chat.completion",
  "model": "openai/gpt-oss-20b",
  "choices": [{
    "message": { "content": "नमस्ते! …" }
  }],
  "usage": {
    "prompt_tokens": 1200,
    "completion_tokens": 350
  },
  "m82": {
    "request_id": "req_01M4M8TAVWH17EPD370FW04AGJ",
    "cost_npr": "0.009037829745"
  }}

Every response has an x-request-id header. Normal responses also carry m82.request_id and the exact m82.cost_npr.

Paying for an API from Nepal

The other way takes a dollar card.

The typical way

  • A dollar card or international card
  • Billed in US dollars
  • Apply for the card, submit documents, load it
  • Card loading charge and an exchange spread
  • A dollar price your bank converts later
  • When credit runs out: depends on the provider

With M82 Router

  • eSewa
  • Billed in Nepali rupees
  • Top up and create a key
  • No card, so no card fees
  • An exact rupee amount on every response
  • Prepaid: requests need credit before they run

Card fees and limits vary by bank.

Get started

Your first request is three steps away.

  1. 1Top up with eSewa
  2. 2Create an API key
  3. 3Change your base URL

Questions

Before you top up.

Do I need a dollar card?

No. You add credit in rupees from eSewa, and you never enter a card on M82 Router.

How is a request priced?

Per token, at the model's current rupee price per million tokens for input, cached input and output. M82 Router prices the whole request once and rounds to 12 decimal places. The response shows the exact amount, and the same request ID appears in your Usage page and ledger.

What happens when my balance runs low?

Before each request, M82 Router holds credit for an estimate of its cost. If your available balance can't cover that hold, the request is refused with a 402 error and nothing is sent to the model. When it settles, you're charged once for what it actually used and any unused hold is released.

What if a request fails or a stream is cut off?

Keep the request ID. If the model provider can't take the request, the hold is released. If a reply is interrupted, M82 Router settles against the usage the provider reported, and reviews the request when that usage is uncertain. It never guesses. Check your Usage page before retrying.

Do you store my prompts?

No. M82 Router doesn't store or log the content of prompts or completions. It keeps request metadata (model, token counts, cost, timing and the request ID) so you can reconcile every charge.

Who processes my requests?

The model providers that host each model. Your request is sent to them to generate the reply, and their data policies apply while they process it.

Can prices change?

Yes, but never backwards. Each request uses the price that was active when its credit was reserved.

What are the rate limits?

Each API key gets 120 requests and about 10,00,000 tokens a minute, with 8 running at once. Your account as a whole gets 300 requests, about 20,00,000 tokens and 20 at once, so extra keys don't raise it. A request over the limit gets a 429 with a Retry-After header.

Can I get a refund?

Refunds are reviewed by hand during the beta. Contact support with your eSewa payment ID. Write to contact@m82labs.dev.

Which tools work with M82 Router?

Any client that lets you set an OpenAI-compatible base URL and key. We test curl and OpenCode end to end; OpenAI's Python and Node SDKs are next on our list.