Get API key

Uncensored LLM API Documentation

This quickstart guide shows how to integrate our uncensored llm api into your application using standard OpenAI-compatible SDKs. You will learn to send requests, handle streaming, and manage constraints without moral filters.

per 1M input tokens
$0.25
Output tokens / 1M
$1.00
token context
64,000
trial credit
$0.50
requests per minute
300

Base URL and Authentication

To use the uncensored llm api, point your client to https://api.uncensoredllmapi.com/v1. Authentication relies on a single API key tied to your account. Generate this key on the Get API key page after signing up with an email and password. Pass the key via the Authorization: Bearer <key> header in every request. If the key is invalid or expired, you receive a 401 error. Note that regenerating the key revokes the old one immediately.

First Request

Send a POST request to /v1/chat/completions with the model ID uncensored. This model is an open-weight LLM tuned for unrestricted responses to lawful adult content. It does not use the GPT, Claude, or Gemini architectures. Include your message history in the messages array. The context window supports 64,000 tokens total for prompt and completion combined.

curl https://api.uncensoredllmapi.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Python SDK Integration

Use the official OpenAI Python SDK to interact with the API. Set the base_url to our endpoint and provide your API key. The client automatically handles JSON serialization and retry logic. This approach works for simple text generation or complex roleplay scenarios where content filters are not desired. Ensure your SDK version is recent enough to support the streaming parameters if needed.

from openai import OpenAI

client = OpenAI(base_url="https://api.uncensoredllmapi.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Node SDK Usage

For JavaScript or TypeScript projects, initialize the OpenAI client with the custom base URL. The Node SDK behaves identically to the Python version regarding authentication and request structure. You can send standard chat completions or function calls. The uncensored model responds to standard tool definitions without blocking based on topic sensitivity. Monitor your usage dashboard to track credit consumption against the prepaid balance.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.uncensoredllmapi.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Streaming Responses

Enable streaming by setting stream: true in your request. The server returns a Server-Sent Events (SSE) stream containing partial token deliveries. This reduces perceived latency for chatbot interfaces. Each chunk contains a delta of the generated text. Process these chunks to update your UI in real-time. The uncensored AI API maintains the same token limits and pricing for streamed responses as for synchronous requests.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Limits, Errors, and Context

Rate limits are strict: 300 requests per minute per key. If you exceed this, you receive a 429 error. The request body must not exceed 8 MB. If your account has no credit, requests fail with a 402 error. The context window is hard-capped at 64,000 tokens. If you send too many tokens, the API returns an error. There are no SLAs or uptime guarantees. Content filtering is minimal, except for sexual content involving minors, which is always blocked.

Questions and answers

Does regenerating my API key reset the rate limit?

No. The rate limit of 300 requests per minute applies to your account, not just the key string. Generating a new key revokes the old one but does not give you a fresh limit.

Is this API suitable for high-volume production?

It depends on your needs. The API lacks SLAs, backups, or multi-region deployment. It is best for apps that prioritize uncensored content over enterprise-grade uptime guarantees. Monitor your token usage closely as credits do not roll over indefinitely.

What happens if I run out of credits?

Requests will return a 402 error. You must top up your account to resume service. Credits are prepaid and do not expire, but you cannot make requests until funds are available.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.