Base URL and Authentication
To use the uncensored llm api, point your client to https://api.uncensoredllmapi.com/v1. Authentication relies on a single API key tied to your account. Generate this key on the Get API key page after signing up with an email and password. Pass the key via the Authorization: Bearer <key> header in every request. If the key is invalid or expired, you receive a 401 error. Note that regenerating the key revokes the old one immediately.
First Request
Send a POST request to /v1/chat/completions with the model ID uncensored. This model is an open-weight LLM tuned for unrestricted responses to lawful adult content. It does not use the GPT, Claude, or Gemini architectures. Include your message history in the messages array. The context window supports 64,000 tokens total for prompt and completion combined.
curl https://api.uncensoredllmapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Python SDK Integration
Use the official OpenAI Python SDK to interact with the API. Set the base_url to our endpoint and provide your API key. The client automatically handles JSON serialization and retry logic. This approach works for simple text generation or complex roleplay scenarios where content filters are not desired. Ensure your SDK version is recent enough to support the streaming parameters if needed.
from openai import OpenAI
client = OpenAI(base_url="https://api.uncensoredllmapi.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Node SDK Usage
For JavaScript or TypeScript projects, initialize the OpenAI client with the custom base URL. The Node SDK behaves identically to the Python version regarding authentication and request structure. You can send standard chat completions or function calls. The uncensored model responds to standard tool definitions without blocking based on topic sensitivity. Monitor your usage dashboard to track credit consumption against the prepaid balance.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.uncensoredllmapi.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Streaming Responses
Enable streaming by setting stream: true in your request. The server returns a Server-Sent Events (SSE) stream containing partial token deliveries. This reduces perceived latency for chatbot interfaces. Each chunk contains a delta of the generated text. Process these chunks to update your UI in real-time. The uncensored AI API maintains the same token limits and pricing for streamed responses as for synchronous requests.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Limits, Errors, and Context
Rate limits are strict: 300 requests per minute per key. If you exceed this, you receive a 429 error. The request body must not exceed 8 MB. If your account has no credit, requests fail with a 402 error. The context window is hard-capped at 64,000 tokens. If you send too many tokens, the API returns an error. There are no SLAs or uptime guarantees. Content filtering is minimal, except for sexual content involving minors, which is always blocked.