logoPofano

Chat Completions

Create a chat completion with OpenAI-compatible format.

POST /v1/chat/completions

The primary endpoint for conversational AI. Send a series of messages and receive a model-generated response. Supports streaming, function calling, and multimodal inputs.

Parameters#

ParameterTypeRequiredDescription
modelstringYesThe model ID to use (e.g., gpt-4o, claude-3-5-sonnet-20241022).
messagesarrayYesA list of message objects representing the conversation history.
temperaturenumberNoSampling temperature between 0 and 2. Higher values produce more random outputs.
max_tokensintegerNoThe maximum number of tokens to generate.
top_pnumberNoNucleus sampling parameter.
frequency_penaltynumberNoPenalty for frequent tokens, between -2.0 and 2.0.
presence_penaltynumberNoPenalty for new tokens, between -2.0 and 2.0.
stopstring/arrayNoSequences where the API will stop generating.
streambooleanNoIf true, the response is streamed as server-sent events.
toolsarrayNoA list of function definitions for function calling.

Message Object#

Each message in the messages array has the following structure:

{
  "role": "user",
  "content": "Hello, how are you?"
}
FieldTypeRequiredDescription
rolestringYesOne of system, user, assistant, or tool.
contentstring/arrayYesThe message content. Can be a string or an array of content parts for multimodal inputs.

Example Requests#

cURL#

curl https://ai.pofano.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
    "model": "gpt-4o",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "What is the capital of France?"}
    ],
    "temperature": 0.7,
    "max_tokens": 256
  }'

Python#

import requests

response = requests.post(
    "https://ai.pofano.com/v1/chat/completions",
    headers={
        "Authorization": "Bearer YOUR_API_KEY",
        "Content-Type": "application/json"
    },
    json={
        "model": "gpt-4o",
        "messages": [
            {"role": "system", "content": "You are a helpful assistant."},
            {"role": "user", "content": "What is the capital of France?"}
        ],
        "temperature": 0.7,
        "max_tokens": 256
    }
)

print(response.json())

JavaScript#

const response = await fetch("https://ai.pofano.com/v1/chat/completions", {
  method: "POST",
  headers: {
    "Authorization": "Bearer YOUR_API_KEY",
    "Content-Type": "application/json"
  },
  body: JSON.stringify({
    model: "gpt-4o",
    messages: [
      { role: "system", content: "You are a helpful assistant." },
      { role: "user", content: "What is the capital of France?" }
    ],
    temperature: 0.7,
    max_tokens: 256
  })
});

const data = await response.json();
console.log(data);

Response#

{
  "id": "chatcmpl-abc123",
  "object": "chat.completion",
  "created": 1677652288,
  "model": "gpt-4o",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "The capital of France is Paris."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 25,
    "completion_tokens": 8,
    "total_tokens": 33
  }
}

Streaming#

To receive responses incrementally, set stream to true. The response will be sent as server-sent events (SSE):

curl https://ai.pofano.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Tell me a story."}],
    "stream": true
  }'

Each chunk is a JSON object prefixed with data: :

data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}

data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Once"},"finish_reason":null}]}

...

data: [DONE]

On this page