Chat Completions
Create a chat completion with OpenAI-compatible format.
POST /v1/chat/completions
The primary endpoint for conversational AI. Send a series of messages and receive a model-generated response. Supports streaming, function calling, and multimodal inputs.
Parameters#
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | The model ID to use (e.g., gpt-4o, claude-3-5-sonnet-20241022). |
messages | array | Yes | A list of message objects representing the conversation history. |
temperature | number | No | Sampling temperature between 0 and 2. Higher values produce more random outputs. |
max_tokens | integer | No | The maximum number of tokens to generate. |
top_p | number | No | Nucleus sampling parameter. |
frequency_penalty | number | No | Penalty for frequent tokens, between -2.0 and 2.0. |
presence_penalty | number | No | Penalty for new tokens, between -2.0 and 2.0. |
stop | string/array | No | Sequences where the API will stop generating. |
stream | boolean | No | If true, the response is streamed as server-sent events. |
tools | array | No | A list of function definitions for function calling. |
Message Object#
Each message in the messages array has the following structure:
{
"role": "user",
"content": "Hello, how are you?"
}| Field | Type | Required | Description |
|---|---|---|---|
role | string | Yes | One of system, user, assistant, or tool. |
content | string/array | Yes | The message content. Can be a string or an array of content parts for multimodal inputs. |
Example Requests#
cURL#
curl https://ai.pofano.com/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "gpt-4o",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is the capital of France?"}
],
"temperature": 0.7,
"max_tokens": 256
}'Python#
import requests
response = requests.post(
"https://ai.pofano.com/v1/chat/completions",
headers={
"Authorization": "Bearer YOUR_API_KEY",
"Content-Type": "application/json"
},
json={
"model": "gpt-4o",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is the capital of France?"}
],
"temperature": 0.7,
"max_tokens": 256
}
)
print(response.json())JavaScript#
const response = await fetch("https://ai.pofano.com/v1/chat/completions", {
method: "POST",
headers: {
"Authorization": "Bearer YOUR_API_KEY",
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "gpt-4o",
messages: [
{ role: "system", content: "You are a helpful assistant." },
{ role: "user", content: "What is the capital of France?" }
],
temperature: 0.7,
max_tokens: 256
})
});
const data = await response.json();
console.log(data);Response#
{
"id": "chatcmpl-abc123",
"object": "chat.completion",
"created": 1677652288,
"model": "gpt-4o",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "The capital of France is Paris."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 25,
"completion_tokens": 8,
"total_tokens": 33
}
}Streaming#
To receive responses incrementally, set stream to true. The response will be sent as server-sent events (SSE):
curl https://ai.pofano.com/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Tell me a story."}],
"stream": true
}'Each chunk is a JSON object prefixed with data: :
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Once"},"finish_reason":null}]}
...
data: [DONE]
Pofano