Every API Call Starts Fresh

Ma Mahalakshmi V Updated 19 Sep 2026
1 min read

The Core Fact: Every API Call Starts From Zero

When you send a request to client.responses.create(), the model that answers you has no idea it has ever spoken to you before. It doesn't remember the question you asked ten seconds ago. It doesn't remember your name, even if you introduced yourself in the previous call. It doesn't even know your last call happened, unless you explicitly tell it what happened by including that information in the current request.

This is one of the most important mental models to build correctly when you start working with the OpenAI API, because it explains almost every "why doesn't the chatbot remember what I said" bug a beginner runs into. The model isn't broken. It isn't being forgetful. It is doing exactly what a stateless API is designed to do: process the one request it was given, in isolation, and return a response.

Let's prove this to ourselves with code before going any further into the theory, because seeing the behavior directly makes the rest of this lesson click much faster.

Setup

You'll need the OpenAI Python SDK and an API key.

Installation

pip install openai python-dotenv

.env

OPENAI_API_KEY=your_api_key_here

Never commit this file or hard-code your key directly into a script. Load it from the environment instead:

import os
from dotenv import load_dotenv
from openai import OpenAI

load_dotenv()
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

In production, set OPENAI_API_KEY through your hosting platform's secrets manager (Render, Railway, AWS Secrets Manager, a Kubernetes secret, and so on) instead of shipping a .env file with your deployment at all. The .env file is a local-development convenience only.

Proving Statelessness With Two Calls

Here's the experiment. We'll tell the model our name in one call, then ask it to recall that name in a completely separate call.

0 Comments

Reviewed before they appear

No comments yet.

OpenAI SDK
Introduction to the OpenAI SDK Setting Up Python Creating an API Key Your First Call — client.responses.create() and response.out Understanding Billing, Credits, and What a Request Costs Why Responses Replaced Chat Completions Anatomy of a Request: model, input, and instructions Anatomy of a Response: The Typed output Array, Not Just Text Roles: User, Assistant, and Developer/System Choosing a Model, and Reading the Models Page Instead of Mem Instructions vs. Input Writing Prompts That Get Consistent Results Few-Shot Examples Reasoning Models and the reasoning Parameter Debugging a Prompt That Misbehaves Why Streaming Matters for User Experience stream=True and Iterating Over Events Handling the Event Types You Actually Care About Background Mode for Long-Running Jobs Project — Add Live Streaming to Your Chatbot The Problem With Parsing Free Text JSON Schema and Strict Mode Pydantic Models With the SDK's Parse Helpers Handling Refusals and Validation Failures Project — A Resume-to-JSON Extractor Working With input_image input_file, PDFs, and the Files API Image Generation Speech-to-Text and Text-to-Speech Project: A PDF Question-Answering Script What Function Calling Is Defining a Tool Schema The Full Loop Multiple Tools Errors, Timeouts, and Untrusted Arguments Project: A Weather Assistant Web Search File Search and Vector Stores Code Interpreter Remote MCP Servers and Connectors Project: A Research Assistant What an Embedding Is, Without the Maths Generating and Storing Embeddings Similarity Search From Scratch Hosted Vector Stores vs. Rolling Your Own A Small RAG App Over a Folder of Notes Agents vs. a Single API Call — When You Need One pip install openai Giving Agents Tools Handoffs and Multi-Agent Triage Guardrails and Approvals Tracing and Observing What Your Agent Did A Multi-Agent Support Desk Error Codes and What Each One Means Retries, Timeouts, and Backoff Rate Limits and Spend Limits Prompt Caching and Cost Optimisation The Batch API for Bulk Work Async Clients and Concurrency Moderation and Safety Best Practices Designing the App Backend With FastAPI Streaming to a Simple Frontend Deploying and a Cost/Safety Checklist Why Web Search Is Useful for Current Information Using the Web Search Tool with the Responses API Configuring Search Behavior for Application Use Cases Understanding Citations and Source Attribution Where to Go Next Building a Research Assistant with Web Search Combining Web Search with Structured Outputs Handling Conflicting or Low-Quality Web Sources
Ask about this post
AI Ask about this post

Ask questions about Every API Call Starts Fresh and get answers drawn from it.

Signed-in readers only.