Accept Google Application Default Credentials (authorized_user) in the
Vertex apiKey field as an alternative to Service Account JSON, for orgs
that block SA key creation. Refreshes a Bearer token via the existing
refreshGoogleToken flow and requires a project_id (quota_project_id or
providerSpecificData.projectId). SA JSON and raw key flows unchanged.
Co-authored-by: Cursor <cursoragent@cursor.com>
Third-party Anthropic-compatible gateways that require Authorization: Bearer
(in addition to x-api-key) returned 401 missing_api_key on the forward path.
For non-official upstreams, also send Bearer <apiKey> alongside x-api-key.
Official api.anthropic.com behavior is unchanged.
Co-authored-by: Cursor <cursoragent@cursor.com>
AWS OIDC IDC/Builder-ID tokens omit profileArn, so CodeWhisperer calls
return 403 "User is not authorized". Resolve it natively via the
ListAvailableProfiles API instead of reading Kiro IDE profile.json.
- providers.js: add fetchKiroProfileArn() and resolve on poll (new logins)
- tokenRefresh.js: backfill profileArn on refresh so existing IDC
connections self-heal without re-login
Co-authored-by: Cursor <cursoragent@cursor.com>
Wire the Kiro provider into BaseExecutor baseUrls fallback so a request
advances across the three CodeWhisperer surfaces (runtime kiro.dev,
codewhisperer, q) on 429 and network/5xx errors.
- providers.js: add baseUrls[] (newest endpoint first); keep baseUrl as default
- kiro.js: replace the hand-rolled single-endpoint loop with super.execute()
delegation plus EventStream to SSE transform on success
Note: the three hosts are alternate DNS surfaces of one regional service. AWS
throttles per authenticated identity (token plus profileArn), not per hostname,
so this is edge-level failover, not extra 429 quota.
Co-authored-by: thienpv <pvtcwd@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
CommandCode upstream only speaks NDJSON streaming; the executor always
wraps it as OpenAI SSE. Force body.stream=true so non-stream client
requests still produce a parseable upstream stream (router aggregates
SSE→JSON downstream).
Co-authored-by: Cursor <cursoragent@cursor.com>
convertOpenAIToolChoice mapped {type:"function"} verbatim and cloakClaudeTools
left tool_choice.name unsuffixed, both rejected by Claude on the cc/ path.
Map forced-function to {type:"tool",name}, allowlist Claude-valid types, and
suffix tool_choice.name when it targets a renamed client tool.
Fixes#1592
Co-authored-by: Cursor <cursoragent@cursor.com>
Add wenyan-lite/wenyan/wenyan-ultra levels for max token compression,
sync SHARED_EXAMPLES/AUTO_CLARITY/PERSISTENCE across all levels, and
expose 3 wenyan buttons in endpoint settings UI.
Co-authored-by: Cursor <cursoragent@cursor.com>
MiniMax requires reasoning_content echoed back on assistant messages in
multi-turn/tool-call conversations. Add minimax and minimax-cn to
PROVIDER_RULES (scope all), same fix as DeepSeek (#1543).
Co-authored-by: Cursor <cursoragent@cursor.com>
Kiro requires a non-empty currentMessage tools array whenever history
references any tool use, else returns "Improperly formed request" (400).
Clients trip this by omitting tools on follow-ups after client-side
compaction.
- flattenToolInteractions(): no client tools -> collapse tool_use/result
to text so the "tools required" rule never fires
- reconcileOrphanedToolResults(): client tools -> salvage orphaned
results as text, keep matched ones, guard co-located tools array
- safeJSONParse(): guard tool-call argument parsing against bad JSON
- merge consecutive user userInputMessageContext; null-guard
currentMessage for assistant-only input
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Add shared OAuth credential lifecycle manager with provider-aware refresh
decisions. Implement CodexExecutor.refreshCredentials so 401/403 retry
refresh works for Codex, track lastRefreshAt and refresh before the
upstream stale-token window, preserve omitted idToken, and add
per-connection single-flight refresh to avoid refresh-token rotation races.
Merged from PR #1664.
Co-authored-by: Cursor <cursoragent@cursor.com>
- Add qmodel_latest to QODER_MODEL_MAP
- Expose qmodel_latest in static Qoder provider catalog (qd)
- Generalize executor comment so model set does not go stale
- Add unit coverage for the new model key + catalog
Closes#1638
Co-authored-by: Cursor <cursoragent@cursor.com>
- Add gemini-3.5-flash-extra-low across CLI menu, provider models, usage, pricing
- Add MITM synonyms (high/medium/extra-low) and split pattern so Low no longer falls through to Medium
- Strip models/ prefix in getMappedModel for AG public name normalization
Co-authored-by: Cursor <cursoragent@cursor.com>
Add mimo-v2.5-pro-claude alias routing to the Xiaomi TokenPlan Anthropic-compatible
/anthropic/v1/messages endpoint. Logic lives in a dedicated XiaomiTokenplanExecutor
(config-driven via targetFormat) instead of the shared DefaultExecutor.
Co-authored-by: Cursor <cursoragent@cursor.com>
Route image model tests to /api/v1/images/generations and STT to
/api/v1/audio/transcriptions instead of forcing all non-embedding
models through chat completions. Adds kind-aware pingModelByKind,
hf->huggingface alias, and silent WAV sample for STT reachability.
Scoped to dashboard/internal model testing only; runtime inference
routing is unchanged.
Author: yicone <yicone@gmail.com>
Closes#1628
## Features
- Add Qoder provider: device-flow OAuth, COSY signing, WAF-bypass body encoding, live model catalog, dashboard quota tracker, 11 models (#1372)
- Add new models: Claude Opus 4.8 (Claude Code), GPT 5.4 Mini (Codex)
## Fixes
- DeepSeek thinking mode: echo `reasoning_content` back on follow-up/tool-call turns so OpenCode-free and custom providers no longer 400 with "reasoning_content must be passed back" (#1543)
- Reasoning injector: match deepseek/kimi model ids case-insensitively (covers custom providers using capitalized model names)
- OpenCode suggested-models: include free models without the `-free` suffix, e.g. `big-pickle` (#1535)
## Improvements
- Codex: trim sunset models, keep gpt-5.5 / gpt-5.4 / gpt-5.3-codex family, add gpt-5.4-mini
- volcengine-ark: refresh model list (add DeepSeek-V4-Flash/Pro, drop EOL entries)
- Lower stream stall timeout 35s → 30s for faster hang detection
Wire Qoder credits into the Quota Tracker card grid:
- Add `qoder` to USAGE_SUPPORTED_PROVIDERS so the connection passes the
isUsageEligible filter at /api/providers/client and shows up in
providerOptions on the dashboard.
- Reshape getQoderUsage so quota records (user, organization) live under
`quotas` and scalar metadata (totalUsagePercentage, isQuotaExceeded,
expiresAt) are siblings — the parser used to walk Object.entries(quotas)
and would have rendered `totalUsagePercentage: 0.42` as a "0/0" row.
- Surface Qoder's expiresAt as resetAt on each quota record so the card
shows when credits reset.
- Add a parser branch in ProviderLimits/utils.js: rename internal keys
(user → "Personal", organization → "Organization"), drop empty org
buckets so personal accounts don't render a misleading "0/0 Organization"
row, and forward remaining/unit so the QuotaProgressBar can use them.
- Add Qoder's brand color (#EC4899) to ProviderLimitCard's color map.
42 tests still pass; build clean.
Adds 18 new tests covering the bugs fixed in the previous commit so they
can't silently regress:
- parseExpiry (7 tests): numeric ms-epoch input, numeric strings handled
before Date.parse so "1700000000" doesn't get year-interpreted, RFC3339
strings, expires_in:0 honored as already-expired, 30-day fallback only
when both inputs are missing/invalid
- normalizeMessages (4 tests): system hoisting, multipart text flatten,
multiple system joining, empty input
- wrapQoderSSE (6 tests): the fixed cases — trailing partial line drained
in flush(), no chunks forwarded after [DONE], embedded newlines stripped
from inner body, error envelope produces error chunk + [DONE], non-ok
responses returned unchanged
- expose parseExpiry from auth.js, expose normalizeMessages/wrapQoderSSE
via __test__ from the executor (internals only — not part of the public
API). Marked with comment so the surface is intentional.
42 tests total (24 original + 18 new). Build still clean.
Correctness:
- testUtils: drop checkExpiry so the userinfo URL probe actually runs (revoked
tokens used to look "active" until local 30-day expiry passed)
- auth.parseExpiry: handle numeric expiresAt, swap parseInt before Date.parse
so "2026" doesn't get interpreted as year-2026, treat expires_in:0 as
already-expired instead of fabricating a 30-day default
- providers.mapTokens: synthesize email from userId when fetchUserInfo fails
so OAuth dedup works (re-logins no longer accumulate "Account N" rows)
SSE wrapper:
- wrapQoderSSE: add !doneEmitted guard on success branch (chunks could leak
past [DONE] when an error envelope shared a TCP packet with a valid one)
- flush(): finalize TextDecoder + drain trailing buffer so the chunk carrying
finish_reason is delivered when upstream closes without a final \n
- sanitize literal \n inside inner OpenAI body so SSE framing stays intact
Robustness:
- executor: wrap buildCosyHeaders in try/catch so a missing accessToken
returns 401 (re-auth) instead of bubbling as 500
- executor: short-circuit on missing accessToken before signing
- executor: plumb proxyOptions/signal through buildQoderRequestBody so
proxy-only networks can fetch the model_config catalog
- qoderModels: dedupe concurrent first-time misses with an in-flight Promise
map (parallel chat windows now do 1 upstream fetch instead of N)
- qoderModels: check signal.aborted before addEventListener so a pre-aborted
parent signal cancels the inner fetch immediately
- auth: AbortController + 15s timeout on pollDeviceToken / fetchUserInfo to
prevent hung sockets when openapi.qoder.sh stalls mid-response
UX:
- OAuthModal: derive polling deadline from device-code expires_in (qoder
publishes 300s; the previous fixed 120s caused timeouts when users took
more than 2 minutes on the consent page)
Cleanup:
- delete src/lib/oauth/services/qoder.js — referenced removed config fields
(clientId/clientSecret/tokenUrl/authorizeUrl) and was re-exported from
services/index.js, so any future caller would TypeError on first use
GitHub Copilot's /responses endpoint only serves OpenAI (gpt/codex)
models. gemini-3.1-pro-preview was failing on /chat/completions with a
"not supported" error, getting cached as a codex model, then escalated
to /responses where it 400s with "does not support Responses API".
Add GithubExecutor.supportsResponsesEndpoint() and gate both the cached
/responses route and the 400-fallback on it, so Gemini/Claude always
stay on /chat/completions and the real upstream error surfaces.
Adds tests/unit/github-responses-routing.test.js (5 tests).
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
## Fixes
- Codex: auto-retry when upstream drops mid-stream (no more hangs)
- Codex: fix random 400/404 errors, tool-calling failures, and unstable prompt cache
- MITM: support Antigravity 2.x (updated IDE version detection and DNS/cert flow)
- Sanitize Read tool args to prevent retry loops from non-Anthropic models (#1144)
- Implement json_schema fallback for OpenAI-compatible providers without native Structured Output (#1343)
- Strip empty Read pages argument in OpenAI-to-Claude translator (#1354)
- Forward Gemini output dimensions for embeddings (#1366)
- Resolve setState-in-effect errors in dashboard components (#1362)
- Gemini CLI: reuse stored OAuth project IDs for quota checks and show clearer setup guidance when the project is missing (#1271, #1428)
When json_schema response_format is sent to models that do not natively
support Structured Output (e.g. DeepSeek/Ollama/local LLMs via
openai-compatible-* providers), they often return empty or malformed
content.
Detect response_format.type === "json_schema", inject the schema into the
system prompt, and downgrade response_format to json_object so Structured
Output works transparently across openai-compatible providers.
Gated to provider.startsWith("openai-compatible-") so providers with
native Structured Output support are not downgraded.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: sanitize Read tool args to prevent retry loops from non-Anthropic models
* fix: sanitize invalid Read pages from tool args
Non-Anthropic models sometimes emit optional Read args like pages: "" for
non-PDF files, which Claude Code rejects before the tool runs. Drop invalid
pages values, keep valid PDF page ranges, and coerce numeric string bounds
before clamping limit/offset.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Some OpenAI-compatible providers emit optional string tool parameters
as empty strings (e.g. pages: "") instead of omitting them. Claude
Code rejects pages: "" as invalid, breaking the Read tool for
non-PDF files routed through 9router.
Add sanitizeToolArguments() that parses tool-call arguments and
removes known optional empty-string fields before emitting
input_json_delta back to Claude format. Currently handles the
Read tool pages field specifically.
Includes regression test.
Fixes#1278
Co-authored-by: JoJo <noreply@github.com>
## Features
- Xiaomi MiMo Token Plan: region selector (Singapore / China / Europe) — keys are cluster-specific
- Antigravity: risk confirmation dialog before first connection
- Gemini CLI: surface upstream retry delay on 429 errors
## Fixes
- MITM: cannot kill process on macOS under sudo (lsof not found in PATH)
- Stream: false-positive stall timeout on Claude reasoning / Kiro responses
- Tunnel: cannot re-enable after disable (stuck state)
- Tunnel: cloudflared error messages now include log tail for easier debugging
- Language switcher: applies selected locale immediately on close (#1234)
- Antigravity OAuth: metadata now matches the official client
## Improvements
- Gemini CLI: bump engine to 0.34.0
- Re-hide `qwen` (OAuth EOL) and `iflow` (not ready) providers
Fixes#1226
The Antigravity OAuth flow sent inconsistent client metadata between
the token acquisition phase and the API usage phase. String enum values
(IDE_UNSPECIFIED, PLATFORM_UNSPECIFIED) were used during OAuth token
exchange + loadCodeAssist + onboardUser, while numeric enums (ideType: 9,
platform: <computed>, pluginType: 2) were used in runtime API calls.
Google detected this fingerprint mismatch and blocked 9router accounts.
Replace all string enum occurrences with the correct numeric values:
- src/lib/oauth/constants/oauth.js: loadCodeAssistClientMetadata now
uses getOAuthPlatformEnum() for platform and numeric 9/2 for
ideType/pluginType, matching getOAuthClientMetadata()
- src/lib/oauth/services/antigravity.js: getMetadata() now delegates
to getOAuthClientMetadata() instead of returning hardcoded strings
- src/lib/oauth/providers.js: postExchange metadata now uses
getOAuthClientMetadata() instead of inline string enums
- open-sse/services/usage.js: getGeminiSubscriptionInfo body now uses
CLIENT_METADATA (already imported from appConstants.js) instead of
inline string enums