You can now request higher scheduling priority per request by setting service_tier: "priority" on text inference endpoints (Chat Completions and Responses). The response's service_tier field reports the tier actually applied, and priority rates are billed only when priority is used. For more details, see the Priority Processing docs.