The Gemini Batch API is designed to process large volumes of requests asynchronously at 50% of the standard cost. The target turnaround time is 24 hours, but in majority of cases, it is much quicker.
Use Batch API for large-scale, non-urgent tasks such as data pre-processing or running evaluations where an immediate response is not required.
Creating a batch job
You have two ways to submit your requests in Batch API:
- Inline requests: A list of
GenerateContentRequestobjects directly included in your batch creation request. This is suitable for smaller batches that keep the total request size under 20MB. The output returned from the model is a list ofinlineResponseobjects. - Input file: A JSON Lines (JSONL)
file where each line contains a complete
GenerateContentRequestobject. This method is recommended for larger requests. The output returned from the model is a JSONL file where each line is either aGenerateContentResponseor a status object.
Inline requests
For a small number of requests, you can directly embed the
GenerateContentRequest objects
within your BatchGenerateContentRequest. The
following example calls the
BatchGenerateContent
method with inline requests:
Python
from google import genai
from google.genai import types
client = genai.Client()
# A list of dictionaries, where each is a GenerateContentRequest
inline_requests = [
{
'contents': [{
'parts': [{'text': 'Tell me a one-sentence joke.'}],
'role': 'user'
}]
},
{
'contents': [{
'parts': [{'text': 'Why is the sky blue?'}],
'role': 'user'
}]
}
]
inline_batch_job = client.batches.create(
model="gemini-3.8-flash",
src=inline_requests,
config={
'display_name': "inlined-requests-job-1",
},
)
print(f"Created batch job: {inline_batch_job.name}")
JavaScript
import {GoogleGenAI} from '@google/genai';
const ai = new GoogleGenAI({});
const inlinedRequests = [
{
contents: