API Reference¶
The FIRST Inference Gateway provides an OpenAI-compatible API.
Base URL¶
Authentication¶
All requests require a Globus access token in the Authorization header:
Endpoints¶
Chat Completions¶
Completions¶
Model Metadata¶
The response contains only models the authenticated caller is authorized to use. Every model includes its identifier, cluster, and framework. Deployments may also expose explicitly allowlisted public metadata from the endpoint fixture, such as a display name, description, and a versioned capabilities object:
[
{
"id": "example-model",
"object": "model",
"cluster": "example-cluster",
"framework": "api",
"display_name": "Example Model",
"description": "Example model served through FIRST",
"capabilities": {
"schema_version": 1,
"api_protocols": ["chat_completions", "responses"],
"context_window_tokens": 131072,
"input_modalities": ["text"],
"streaming": true,
"reasoning": {
"supported": true,
"separate_output": true
},
"tool_calling": {
"supported": true
}
}
}
]
Capability metadata describes the deployed model API. It is intentionally client-neutral and must not contain backend connection details or secrets.
Batch Processing¶
For detailed API documentation, refer to the OpenAI API Reference as FIRST follows the same schema.
Request Parameters¶
See the User Guide for detailed parameter documentation.