AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→

Q: What is the Chat Completions API?

The Chat Completions API is the standard way to talk to an LLM from code. OpenAI introduced it, and it became so widespread that most other providers now offer APIs in the same format. Learn it once, and you can work with almost any model.

The mechanics are simple. From Python you call the OpenAI library, and under the hood it sends an ordinary HTTPS request to OpenAI's servers: a list of messages goes out, the model's next message comes back. Each message has a role: "system" for your instructions, "user" for what the person said, and "assistant" for the model's own earlier replies.

In Python, your first call looks like this:

from openai import OpenAI

client = OpenAI()  # reads your API key from the environment

response = client.chat.completions.create(
    model="gpt-4.1-mini",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Explain RAG in one sentence."}
    ]
)

print(response.choices[0].message.content)

That's the whole thing. No special AI programming, no frameworks.

And here's the part many people don't realize: there's no magic wire into the model. That Python call is a wrapper around a plain HTTPS request. You can make the exact same call from a terminal with curl, no Python involved:

curl https://api.openai.com/v1/chat/completions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4.1-mini",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "Explain RAG in one sentence."}
    ]
  }'

It's a normal web API, the same kind engineers have been calling for years. The AI lives on the server; your side is ordinary engineering.

One thing to understand from the start: the API is stateless. It doesn't remember your previous calls, which is why you send the whole message list every time.

Worth knowing: OpenAI has since released a newer interface called the Responses API, and it will grow in importance. But Chat Completions is the format the industry standardized on, it's everywhere in tutorials, codebases, and job interviews, and it's the right place to start.

← Back to the full FAQ