release notes
grok-voice-transcribe-1.0 is deprecated and reaches end of life on October 2, 2026. All requests to that slug are routed to grok-voice-transcribe-2.0 at the same price, with higher accuracy. See the Speech to Text docs.
ever since claude code, keeping up with ai updates has become its own job
the latest updates to your ai tools, straight from their official changelogs.
grok-voice-transcribe-1.0 is deprecated and reaches end of life on October 2, 2026. All requests to that slug are routed to grok-voice-transcribe-2.0 at the same price, with higher accuracy. See the Speech to Text docs.
You can now send safety_identifier, an opaque end-user identifier assigned by your application, on Chat Completions, the Responses API, deferred chat completions, the Batch API, and the gRPC GetCompletionsRequest. It lets SpaceXAI attribute a policy violation to one of your end users rather than to your API key. The legacy user field is still accepted. See the Security FAQ.
Grok 4.7, SpaceXAI's frontier model for coding, agentic tasks, and knowledge work, is now available on the xAI API as grok-4.7. It has a 500k context window, text and image inputs with text-only output, and no text output limit. Pricing is $2 / $0.50 / $6 per 1M tokens (input / cached input / output) below 200k prompt tokens, and $4 / $1 / $12 above. Reasoning effort supports low, medium, high (default), and xhigh. On the Responses API, grok-4.7 always returns reasoning.encrypted_content, even when include does not list it. It is also served on the US regional endpoint. Grok 4.7 Fast is the same model and costs 2x the standard token rates, or 1.5x for long-context requests. It's available only in Cursor and Grok Build, not on the public xAI API, and is billed through your plan there. See the Grok 4.7 overview and the announcement.
grok-voice-transcribe-2.0 is now available. Use grok-voice-transcribe-1.0 or grok-voice-transcribe-2.0; the default is grok-voice-transcribe-2.0. See the Speech to Text docs.
On November 2, 2026, grok-imagine-image-quality is retired. Requests to the slug will be served by grok-imagine-image-2.0 with quality set to low, with no change to the request or response shape and at a lower per-image price. grok-imagine-image (1.0) is not affected. See the migration guide.
Auto quality. The quality parameter on grok-imagine-image-2.0 now accepts auto, and the default when quality is omitted has moved from medium to auto. Auto currently uses low for image generation and medium for image editing. Images are billed at the quality they are served at. Pass low or medium explicitly to pin a specific quality. See Image Generation. Five reference images. Image editing now accepts up to 5 source images per request (was 3). See Multi-Image Editing. New aspect ratios. Image generation and editing accept 21:9 (cinematic widescreen) and 5:2 (wide banners). See Image Generation.
Grok 4.6, SpaceXAI's frontier model for coding, agentic tasks, and knowledge work, is now available on the xAI API. It has a 500k context window, text and image inputs with text-only output, and no text output limit. Pricing is $2 / $0.50 / $6 per 1M tokens (input / cached input / output) below 200k prompt tokens, and $4 / $1 / $12 above. Reasoning effort supports low, medium, high (default), and xhigh. See the Grok 4.6 overview and the announcement.
grok-imagine-video-1.5 now supports text-to-video, image-to-video, and reference-to-video (including optional preset voices), with native 1080p for T2V and I2V. Text-to-video on this model runs as text-to-image then image-to-video under the hood. See Video Generation, Image-to-Video, and Reference-to-Video.
grok-voice-think-fast-2.0 is now available with Speech to Speech. grok-voice-latest will route to this model starting August 5, 2026. To get started, see the Speech to Speech docs. For more details, see our announcement.
Speech to Text now accepts a vad_threshold parameter (streaming query param and batch multipart field) to tune the voice-activity gate that skips non-speech audio. Lower values transcribe quieter or noisier speech — useful for narrowband telephony — and 0 disables the gate. See the Speech to Text docs.
Grok 4.5 is now available in the API console for EU users. See the Grok 4.5 overview.
Grok 4.5, SpaceXAI's model for coding, agentic tasks, and knowledge work, is now available on the xAI API. Priced at $2 / 1M input tokens and $6 / 1M output tokens, with configurable reasoning effort (low, medium, or high; default high). See the Grok 4.5 overview and the announcement.
You can now request higher scheduling priority per request by setting service_tier: "priority" on text inference endpoints (Chat Completions and Responses). The response's service_tier field reports the tier actually applied, and priority rates are billed only when priority is used. For more details, see the Priority Processing docs.
Public URLs for Files — turn any file in your Files API storage into a permanent, unauthenticated URL that anyone can open, embed, or share. Revocable at any time, or set an auto-expiry between 1 hour and 30 days. See the Public URLs docs. Reference stored files as Imagine inputs — substitute image_file_id, video_file_id, or reference_image_file_ids for URL inputs across every Imagine endpoint, with no need to re-upload bytes or make the file public. See Imagine → Files API Integration. Persist Imagine outputs to Files — set storage_options on any Imagine request to save the generated asset to your Files storage; pair with storage_options.public_url to publish a shareable link in one round trip. See Imagine → Files API Integration.
The streaming Speech to Text API now supports Smart Turn end-of-turn detection. When enabled via the smart_turn query parameter, an ML model predicts whether the speaker has finished their thought at silence boundaries — reducing false endpointing during dictation, number sequences, and mid-sentence pauses. Use smart_turn_timeout to set a maximum silence fallback. For more details, see the Smart Turn docs.
The Context Compaction API is now available. You can shrink long conversations into a shorter context and reuse it in follow-up requests for lower cost, faster time-to-first-token, and sharper responses on long agent loops. For more details, see the Context Compaction docs.
WebSocket Responses API mode is now available. Drive the Responses API over a single, long-lived WebSocket connection for lower end-to-end latency on tool-heavy agent workloads. For more details, see the WebSocket Mode docs.
Web Search now supports explicitly searching for images. Enable enable_image_search to let Grok search directly for relevant images; responses can include returned images as Markdown image embeds. For details, see Enable Image Search.
You can now clone a voice from a short audio clip and use it across the Text-to-Speech and Speech to Speech APIs. Create and manage your voice catalog from the xAI console. For more details, check out the Custom Voices docs and our blog post.
Every API response now includes the exact cost of the request via a cost_in_usd_ticks field in the usage object. Works across chat completions, Responses API, image generation, video generation, and streaming. For more details, see the Cost Tracking docs.
You can now set an expiration policy on uploaded files using expires_after or an explicit expires_at timestamp. Expired files are automatically deleted. For more details, see the Files API docs.
You can now use grok-voice-think-fast-1.0 with the Speech to Speech API. To get started, check out the Speech to Speech docs. For more details, see our blog post.
The xAI Speech to Text API is now generally available. Transcribe audio to text in 25 languages with batch and streaming modes. For more details, check out the Speech to Text docs.
The Text-to-Speech API is now generally available. Generate natural-sounding speech from text with Grok. For more details, check out the Text-to-Speech docs.
The Batch API now supports image generation, image editing, and video generation in addition to chat completions. Both server-side tools and client-side function tools are also now supported in batch requests. Image and video URLs in batch results expire after 1 hour.
You can now create batches by uploading a JSONL file via the Files API. Supports all batch endpoints including chat, image, and video in a single file.
For more details on Grok 4.20 Multi-agent, check out the docs
Batch API is available for all customers. It enables efficient batch processing of multiple requests, providing a better experience for users who need to submit large volumes of requests at once.
Video Generation and a revamped Image Generation are now available.
Grok Speech to Speech API is generally available. Visit Grok Speech to Speech API for guidance on using the API.