AI Coding API Quickstart Guide
Connect your coding agent to an uncensored, OpenAI-compatible endpoint that handles code-heavy prompts without unnecessary content refusals. This guide covers authentication, SDK integration, streaming, and rate limits for the ai coding api.
Authentication Setup
Start by creating an account at Get API key. You only need an email and password; no card is required to start. Upon signup, you receive a single API key immediately. This key authenticates all requests to the base URL: https://api.codingllmapi.com/v1.
Keep your key secure. You can regenerate it at any time from your dashboard, which instantly invalidates the previous key. The model identifier to send in your requests is uncensored. This is an open-weight model running on our own servers, not a proxy for GPT, Claude, or Gemini.
First Request
Send a standard chat completion request using the OpenAI-compatible format. The API accepts messages with a role (user/assistant/system) and content. It returns text output optimized for code generation without filtering lawful adult or controversial topics (except for minors).
- Base URL:
https://api.codingllmapi.com/v1 - Endpoint:
POST /v1/chat/completions - Model:
uncensored
curl https://api.codingllmapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Python SDK Integration
Use the official OpenAI Python SDK. Point the client to our base URL and supply your API key. The SDK handles the request format automatically. This approach works for simple text generation or building more complex agent loops.
Ensure you are using a recent version of the openai package. The client will handle JSON serialization and HTTP headers correctly for our endpoint.
from openai import OpenAI
client = OpenAI(base_url="https://api.codingllmapi.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Node SDK Integration
Use the official openai Node.js package, not the Anthropic or Google-specific SDKs. Configure the base URL to point to our service. This ensures compatibility with the OpenAI-compatible request structure we serve.
Install the package via npm install openai. Initialize the client with your key and the custom base URL. This method is reliable for server-side coding agents that need to generate code snippets or analyze repositories.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.codingllmapi.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Streaming Code Responses
For real-time feedback, enable streaming by setting stream: true. The API returns Server-Sent Events (SSE). Each chunk contains partial token outputs, allowing your agent to display code as it is generated. This reduces perceived latency for long code blocks.
Handle the stream events to accumulate the final response. The model continues generating until it reaches the context limit or the max_tokens parameter.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Limits and Errors
Watch these limits to avoid interruptions. You get 300 requests per minute per key. The maximum request body size is 8 MB. If you exceed the rate limit, you receive a 429 error. If your prepaid credit runs out, you get a 402 error. An invalid or revoked key returns 401.
The context window is 100,000 tokens total (prompt + completion). If your prompt plus generated code exceeds this, the model stops. Plan your input tokens carefully to leave room for the output.
Under the hood: specs
A quick checklist for developers: format, limits, features, billing.
| Parameter | Details |
|---|---|
| Compatibility | OpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key |
| Authentication | Authorization: Bearer YOUR_KEY |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Model ID | uncensored |
| Base URL | https://api.codingllmapi.com/v1 |
| Max output | up to the rest of the 100,000-token window; max_tokens optional (no separate cap) |
| Streaming | Yes — server-sent events; the last chunk carries token usage |
| Max context | 100,000 tokens, input and output combined |
| Sampling parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Structured output | JSON object mode via response_format json_object |
| Tools / tool calls | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Parallel requests | 8 requests at the same time per key |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Max body | 8 MB request body |
| Rate limit | 300 requests per minute per key |
| Free trial | $0.50 for 7 days, no card · Trial key: 2 parallel requests, 60 req/min; full limits (8 and 300) after first top-up |
| Volume bonus | +5% on $50+, +10% on $100+ |
| Price | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Top-up | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Subscription | no monthly fee; paid credit does not expire |
| Billing | pay as you go from prepaid credit; nothing is charged for failed or refused requests |
| Keys | one active key per account; a new key replaces the old one |
| Content policy | adult content allowed; sexual content involving minors is refused |
| Sign-in | Google or e-mail and password |
When a request fails
The type field is stable, the message is for humans. Errors cost nothing.
| Status | Type | Reason |
|---|---|---|
400 | bad_request | malformed request or too long for the context window |
401 | missing_key · invalid_key · key_revoked | check the Authorization header or use your current key |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | refused by the content policy |
404 | not_found | unknown endpoint |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
Is this model the same as GPT-4 or Claude?
No. We serve one uncensored open-weight model on our own GPU servers. It is tuned for code generation and does not use content filters for lawful adult or controversial topics, unlike general-purpose models. It is not a proxy for OpenAI or Anthropic.
How does pricing work?
We use a prepaid credit model. Input tokens cost $0.25 per 1M; output tokens cost $1.00 per 1M. Credits never expire. You can top up with crypto (USDT or USDC), starting at $10. No monthly fees or subscriptions.
What content is blocked?
The only hard limit is sexual content involving minors. All other lawful adult, fictional, or controversial topics are allowed. We do not use prompts for training. Your data stays private to your account.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.
Get API key