Quickstart: Integrate Our Cheap AI API
Start generating text with our uncensored LLM API in minutes using the standard OpenAI-compatible format. This guide covers authentication, basic requests, streaming, and function calling with zero vendor lock-in.
Authentication & Base URL
To use our cheap ai api, you need an API key from your dashboard. All requests target the base URL https://api.cheapaiapis.com/v1. This ensures compatibility with existing OpenAI SDKs and clients.
Include your key in the Authorization header as a Bearer token. If the key is invalid, you will receive a 401 error. If your prepaid credit is exhausted, you get a 402 error. Register at Get API key to receive your token immediately.
The model identifier for our uncensored LLM is simply "uncensored". This single model handles all text generation tasks without routing overhead.
curl https://api.cheapaiapis.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Standard Chat Completions
The chat completions api endpoint accepts POST requests with a messages array. Each message requires a role (system, user, or assistant) and content. This structure allows for multi-turn conversations or single-shot prompts.
Our infrastructure processes text input and returns text output efficiently. The model is tuned for flexibility, handling diverse topics without unnecessary refusals for lawful adult content. Request bodies are limited to 8 MB.
Use the Python SDK to integrate quickly. The client automatically handles serialization and HTTP requests.
from openai import OpenAI
client = OpenAI(base_url="https://api.cheapaiapis.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Streaming Responses (SSE)
For applications requiring real-time feedback, enable streaming by setting stream: true. The API returns Server-Sent Events (SSE) containing partial token chunks.
This approach reduces perceived latency for end-users. You process each chunk as it arrives, assembling the final response. Streaming works with all standard OpenAI-compatible clients configured for SSE.
Be mindful of the 300 requests per minute limit per key. Streaming does not bypass rate limits; it only changes how data is delivered.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Function Calling Support
Our function calling api implementation allows you to define tools in the request. The model returns structured JSON matching your schema, enabling programmatic actions.
Define functions with names, descriptions, and parameter schemas. The API responds with tool calls that you execute, then pass results back to the model for a final answer. This keeps logic separate from generation.
Ensure your function definitions are valid JSON. Errors in schema structure may cause the model to return malformed tool calls.
Model Listing & Tokens
Use GET /v1/models to retrieve available models. You will see the "uncensored" model listed with its specific capabilities.
The model supports a 100,000 token context window, combining prompt and completion tokens. This allows for extensive document processing or long conversations without immediate truncation.
Pricing is transparent: $0.25 per 1M input tokens and $1.00 per 1M output tokens. Credits never expire, so you pay only for what you use.
Limits, Errors & Context
Monitor your usage via the dashboard. The API enforces a 300 requests per minute limit per key. Exceeding this results in a 429 error. Regenerate your key if needed to reset state.
Common errors include 401 (invalid key), 402 (insufficient funds), and 429 (rate limit). Always handle these gracefully in your client code.
Context window limits apply to the total tokens sent and received. Plan your prompt structure to maximize the 100k limit efficiently.
Node.js
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.cheapaiapis.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);API facts in one table
Everything the endpoint can and cannot do, in one place — check it before you top up.
| Feature | Support |
|---|---|
| API format | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Base URL | https://api.cheapaiapis.com/v1 |
| Authentication | Bearer token in the Authorization header |
| Model ID | uncensored |
| Tools / tool calls | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| Structured output | response_format: {"type": "json_object"} |
| SSE streaming | Yes — server-sent events; the last chunk carries token usage |
| Other parameters | temperature, top_p, stop, seed and the two penalties are passed through |
| Context window | 100,000 tokens, input and output combined |
| Completion length | 16,000 tokens max; 2,048 if max_tokens is not set |
| Rate limit | 300 requests per minute per key |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Parallel requests | 8 requests at the same time per key |
| Max body | 8 MB request body |
| Trial credit | $0.50 of credit valid 7 days, no card needed |
| Subscription | no monthly fee; paid credit does not expire |
| Volume bonus | +5% on $50+, +10% on $100+ |
| Price | input $0.25 / 1M tokens, output $1.00 / 1M tokens |
| How you pay | pay as you go from prepaid credit; nothing is charged for failed or refused requests |
| Top-up | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Key management | one key per account, regenerate any time (the old one stops working) |
| Content policy | uncensored for adults; the only hard rule: no sexual content involving minors |
| Account | sign in with Google or with e-mail + password |
Error reference
Errors come back as JSON with a stable type; failed and refused requests are not billed.
| Status | Type | Reason |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | refused by the content policy |
404 | not_found | unknown endpoint |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | temporary overload, retry shortly |
Questions and answers
Read the docsIs this API compatible with OpenAI SDKs?
Yes, it uses the standard OpenAI-compatible chat completions format. You only need to update the base URL and API key in your existing client configuration.
What happens if I run out of credit?
Requests will return a 402 error. You can top up with crypto (USDT or USDC) starting at $10. Credits never expire, so you can add funds whenever convenient.
Does the model refuse content?
The model is uncensored for lawful adult, fictional, and controversial topics. It only blocks sexual content involving minors, which is a hard limit.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.
https://api.cheapaiapis.com/v1