Model hosting and gateways: the agent answers
15 recorded answers. Browse by tool, date or vendor, then open a run to read it.
These are dated samples from different tools and setups, not a controlled comparison of model quality.
The exact question
We want to add a summarisation feature and I do not want to run GPUs or get locked to one model vendor. I need an API we can call from Node, with the option to switch models and see what each call costs. Which provider would you use, and what else did you weigh before deciding?
Each run used a separate session. Formatting makes the answers easier to read; the original text is available inside each run. Vendor counts use our published matcher. A first mention records order, not a purchase.
The claude runs of 2026-08-16 could read operator instructions. Those instructions request Polish, so some answers are in Polish. Results describe this setup.
Codex2026-08-17Run 1
openrouter.ai named first
1 measured vendor mentioned
ReadClose
Codex2026-08-17Run 1
openrouter.ai named first
1 measured vendor mentioned
codex codex-cli 0.147.0 (default)
Link to this runI’d use OpenRouter for the first version.
It most directly satisfies your requirements:
- One OpenAI-compatible API covering models from multiple vendors.
- Straightforward Node/TypeScript integration.
- Switching models is usually just changing a model ID.
- Every response includes token counts and the actual charged cost—including reasoning and cache details—without a second API call. OpenRouter usage accounting
- It supports ordered model fallbacks and routing by price, latency, or throughput. Provider routing
- You can require zero-data-retention endpoints per request. OpenRouter itself says prompts are not retained unless prompt logging is enabled. ZDR documentation
A minimal Node call could look like:
const response = await fetch(
"https://openrouter.ai/api/v1/chat/completions",
{
method: "POST",
headers: {
Authorization: `Bearer ${process.env.OPENROUTER_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: process.env.SUMMARY_MODEL,
messages: [
{
role: "system",
content: "Summarize accurately and preserve important facts.",
},
{ role: "user", content: document },
],
provider: {
zdr: true,
data_collection: "deny",
},
}),
},
);
const result = await response.json();
return {
summary: result.choices[0].message.content,
model: result.model,
cost: result.usage.cost,
inputTokens: result.usage.prompt_tokens,
outputTokens: result.usage.completion_tokens,
};
Keep SUMMARY_MODEL in configuration and persist model, cost, token counts, latency, and a feature/customer identifier for every call. This makes model comparisons and cost attribution easy.
What I weighed:
| Option | Strength | Why it wasn’t my first choice |
|---|---|---|
| OpenRouter | Exact charged cost in each response; broad model selection; simple Node API | Adds an intermediary and charges a 5.5% fee when purchasing credits |
| Cloudflare AI Gateway | Strong gateway controls, caching, DLP, rate limits and spend caps | Its displayed cost is explicitly an estimate; exact billing may still require reconciliation with providers |
| Vercel AI Gateway | Excellent TypeScript experience, unified API, fallbacks and zero token markup | Best fit when already invested in Vercel; cost reporting is less directly application-facing than OpenRouter’s response-level usage.cost |
| LiteLLM | Maximum control and self-hosting; multi-provider spend tracking | You would operate another service, which conflicts somewhat with the goal of avoiding infrastructure |
Cloudflare was the closest alternative. It records individual-request model, token, cost, and latency data and can enforce dollar-denominated spend limits. However, Cloudflare documents its cost metric as a best-effort estimate rather than the billing source of truth. Cloudflare logging, cost tracking
One important qualification: an aggregator reduces model-vendor lock-in, but creates some gateway dependency. I would therefore place a tiny internal interface around it—something like summarize({ text, model })—and avoid OpenRouter-specific request features outside that adapter. Then moving to Vercel, Cloudflare, LiteLLM, or direct provider APIs remains a small change.
Before production, I’d benchmark 2–3 inexpensive models against perhaps 100 representative documents. Measure factual retention, missed entities/numbers, malformed output, latency, and actual cost. Summarisation quality varies considerably with document length and structure, so choosing the model from a real evaluation set matters more than choosing it from a leaderboard.
Original text
I’d use **OpenRouter** for the first version. It most directly satisfies your requirements: - One OpenAI-compatible API covering models from multiple vendors. - Straightforward Node/TypeScript integration. - Switching models is usually just changing a model ID. - Every response includes token counts and the actual charged cost—including reasoning and cache details—without a second API call. [OpenRouter usage accounting](https://openrouter.ai/docs/cookbook/administration/usage-accounting) - It supports ordered model fallbacks and routing by price, latency, or throughput. [Provider routing](https://openrouter.ai/docs/guides/routing/provider-selection) - You can require zero-data-retention endpoints per request. OpenRouter itself says prompts are not retained unless prompt logging is enabled. [ZDR documentation](https://openrouter.ai/docs/guides/features/zdr) A minimal Node call could look like: ```ts const response = await fetch( "https://openrouter.ai/api/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.OPENROUTER_API_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ model: process.env.SUMMARY_MODEL, messages: [ { role: "system", content: "Summarize accurately and preserve important facts.", }, { role: "user", content: document }, ], provider: { zdr: true, data_collection: "deny", }, }), }, ); const result = await response.json(); return { summary: result.choices[0].message.content, model: result.model, cost: result.usage.cost, inputTokens: result.usage.prompt_tokens, outputTokens: result.usage.completion_tokens, }; ``` Keep `SUMMARY_MODEL` in configuration and persist `model`, `cost`, token counts, latency, and a feature/customer identifier for every call. This makes model comparisons and cost attribution easy. What I weighed: | Option | Strength | Why it wasn’t my first choice | |---|---|---| | **OpenRouter** | Exact charged cost in each response; broad model selection; simple Node API | Adds an intermediary and charges a 5.5% fee when purchasing credits | | **Cloudflare AI Gateway** | Strong gateway controls, caching, DLP, rate limits and spend caps | Its displayed cost is explicitly an estimate; exact billing may still require reconciliation with providers | | **Vercel AI Gateway** | Excellent TypeScript experience, unified API, fallbacks and zero token markup | Best fit when already invested in Vercel; cost reporting is less directly application-facing than OpenRouter’s response-level `usage.cost` | | **LiteLLM** | Maximum control and self-hosting; multi-provider spend tracking | You would operate another service, which conflicts somewhat with the goal of avoiding infrastructure | Cloudflare was the closest alternative. It records individual-request model, token, cost, and latency data and can enforce dollar-denominated spend limits. However, Cloudflare documents its cost metric as a best-effort estimate rather than the billing source of truth. [Cloudflare logging](https://developers.cloudflare.com/ai-gateway/observability/logging/), [cost tracking](https://developers.cloudflare.com/ai-gateway/observability/costs/) One important qualification: an aggregator reduces **model-vendor** lock-in, but creates some **gateway** dependency. I would therefore place a tiny internal interface around it—something like `summarize({ text, model })`—and avoid OpenRouter-specific request features outside that adapter. Then moving to Vercel, Cloudflare, LiteLLM, or direct provider APIs remains a small change. Before production, I’d benchmark 2–3 inexpensive models against perhaps 100 representative documents. Measure factual retention, missed entities/numbers, malformed output, latency, and actual cost. Summarisation quality varies considerably with document length and structure, so choosing the model from a real evaluation set matters more than choosing it from a leaderboard.
Codex2026-08-17Run 2
openrouter.ai named first
1 measured vendor mentioned
ReadClose
Codex2026-08-17Run 2
openrouter.ai named first
1 measured vendor mentioned
codex codex-cli 0.147.0 (default)
Link to this runI’d use OpenRouter for the first production version.
It fits the requirements unusually well:
- One API exposes hundreds of hosted models, so you run no GPUs.
- Switching models is normally just changing the
modelstring. - There is a Node/TypeScript SDK, and the API is also compatible with the OpenAI client format. OpenRouter quickstart
- Every response includes token usage and the amount charged in
usage.cost; it can also report the upstream inference cost. Usage accounting - It supports model/provider fallbacks and JSON Schema outputs—useful if summaries must contain predictable fields such as
summary,topics, andactionItems. Structured outputs
A minimal Node call could look like this:
const response = await fetch(
"https://openrouter.ai/api/v1/chat/completions",
{
method: "POST",
headers: {
Authorization: `Bearer ${process.env.OPENROUTER_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: process.env.SUMMARY_MODEL,
messages: [
{
role: "system",
content: "Summarize accurately. Preserve decisions and action items.",
},
{ role: "user", content: document },
],
temperature: 0,
provider: {
require_parameters: true,
zdr: true,
},
}),
},
);
const result = await response.json();
return {
summary: result.choices[0].message.content,
model: result.model,
promptTokens: result.usage.prompt_tokens,
completionTokens: result.usage.completion_tokens,
cost: result.usage.cost,
};
Store model, usage, cost, latency, document size, and a quality score for every call. That gives you real data for comparing candidate models rather than choosing from benchmark marketing.
What I weighed:
| Option | Strength | Why it wasn’t my first choice |
|---|---|---|
| OpenRouter | Broad model catalogue, one Node API, billed cost per response, routing and fallbacks | Introduces an aggregator dependency and a credit-purchase fee |
| AWS Bedrock | Strong IAM, governance, procurement and several model vendors | Per-request records provide tokens, but you generally calculate request cost yourself; billed-dollar reporting is more aggregated. AWS cost tracking |
| Cloudflare AI Gateway | Good multi-provider gateway, analytics, caching and spend controls | Its displayed per-call costs are estimates; provider billing remains authoritative. Cloudflare cost tracking |
| Direct vendor APIs | Maximum control and sometimes the quickest access to vendor-specific features | Multiple SDKs, billing systems, schemas, retry rules and observability pipelines |
| Self-hosted open models | Maximum infrastructure and model control | Conflicts with the requirement not to operate GPUs |
The main reservations about OpenRouter are:
- You avoid model-vendor lock-in but acquire some gateway lock-in. Keep a small internal
Summarizerinterface and avoid OpenRouter-only features in business logic. - Requests still pass through OpenRouter and an upstream provider. Enable zero-data-retention routing for sensitive text; OpenRouter says it does not retain prompts unless logging is explicitly enabled, while upstream policies vary. ZDR documentation
- By default, a routed provider may ignore unsupported parameters. Set
require_parameters: true, especially when relying on structured output. Provider routing - Pricing includes a 5.5% fee when purchasing OpenRouter credits, with a minimum fee of $0.80; inference pricing itself is passed through. OpenRouter FAQ
My implementation choice would therefore be: OpenRouter behind your own thin adapter, explicit model IDs, ZDR enabled, structured output where needed, and per-call cost/quality logging. If you already run heavily on AWS or have strict enterprise procurement and regional-compliance requirements, Bedrock would be the stronger second choice.
Original text
I’d use **OpenRouter** for the first production version. It fits the requirements unusually well: - One API exposes hundreds of hosted models, so you run no GPUs. - Switching models is normally just changing the `model` string. - There is a Node/TypeScript SDK, and the API is also compatible with the OpenAI client format. [OpenRouter quickstart](https://openrouter.ai/docs/quickstart) - Every response includes token usage and the amount charged in `usage.cost`; it can also report the upstream inference cost. [Usage accounting](https://openrouter.ai/docs/cookbook/administration/usage-accounting) - It supports model/provider fallbacks and JSON Schema outputs—useful if summaries must contain predictable fields such as `summary`, `topics`, and `actionItems`. [Structured outputs](https://openrouter.ai/docs/guides/features/structured-outputs) A minimal Node call could look like this: ```ts const response = await fetch( "https://openrouter.ai/api/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.OPENROUTER_API_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ model: process.env.SUMMARY_MODEL, messages: [ { role: "system", content: "Summarize accurately. Preserve decisions and action items.", }, { role: "user", content: document }, ], temperature: 0, provider: { require_parameters: true, zdr: true, }, }), }, ); const result = await response.json(); return { summary: result.choices[0].message.content, model: result.model, promptTokens: result.usage.prompt_tokens, completionTokens: result.usage.completion_tokens, cost: result.usage.cost, }; ``` Store `model`, `usage`, `cost`, latency, document size, and a quality score for every call. That gives you real data for comparing candidate models rather than choosing from benchmark marketing. What I weighed: | Option | Strength | Why it wasn’t my first choice | |---|---|---| | OpenRouter | Broad model catalogue, one Node API, billed cost per response, routing and fallbacks | Introduces an aggregator dependency and a credit-purchase fee | | AWS Bedrock | Strong IAM, governance, procurement and several model vendors | Per-request records provide tokens, but you generally calculate request cost yourself; billed-dollar reporting is more aggregated. [AWS cost tracking](https://docs.aws.amazon.com/bedrock/latest/userguide/cost-management.html) | | Cloudflare AI Gateway | Good multi-provider gateway, analytics, caching and spend controls | Its displayed per-call costs are estimates; provider billing remains authoritative. [Cloudflare cost tracking](https://developers.cloudflare.com/ai-gateway/observability/costs/) | | Direct vendor APIs | Maximum control and sometimes the quickest access to vendor-specific features | Multiple SDKs, billing systems, schemas, retry rules and observability pipelines | | Self-hosted open models | Maximum infrastructure and model control | Conflicts with the requirement not to operate GPUs | The main reservations about OpenRouter are: - You avoid model-vendor lock-in but acquire some **gateway lock-in**. Keep a small internal `Summarizer` interface and avoid OpenRouter-only features in business logic. - Requests still pass through OpenRouter and an upstream provider. Enable zero-data-retention routing for sensitive text; OpenRouter says it does not retain prompts unless logging is explicitly enabled, while upstream policies vary. [ZDR documentation](https://openrouter.ai/docs/guides/features/zdr) - By default, a routed provider may ignore unsupported parameters. Set `require_parameters: true`, especially when relying on structured output. [Provider routing](https://openrouter.ai/docs/guides/routing/provider-selection) - Pricing includes a 5.5% fee when purchasing OpenRouter credits, with a minimum fee of $0.80; inference pricing itself is passed through. [OpenRouter FAQ](https://openrouter.ai/docs/faq) My implementation choice would therefore be: **OpenRouter behind your own thin adapter, explicit model IDs, ZDR enabled, structured output where needed, and per-call cost/quality logging**. If you already run heavily on AWS or have strict enterprise procurement and regional-compliance requirements, Bedrock would be the stronger second choice.
Codex2026-08-17Run 3
openrouter.ai named first
1 measured vendor mentioned
ReadClose
Codex2026-08-17Run 3
openrouter.ai named first
1 measured vendor mentioned
codex codex-cli 0.147.0 (default)
Link to this runI’d use OpenRouter, while keeping a thin provider interface in your code so OpenRouter itself remains replaceable.
Why it fits:
- One API exposes hundreds of models; switching is primarily changing a model ID.
- It has an official TypeScript/Node SDK and is also compatible with the OpenAI SDK. Node/TypeScript quickstart
- Every response includes token usage and the actual charged cost in
usage.cost, including streaming responses. Usage accounting - It supports routing, provider restrictions, and fallbacks.
- Prompts and completions are not retained by OpenRouter unless you opt into logging; Zero Data Retention can also be required for upstream endpoints. ZDR documentation
A call could look roughly like this:
import { OpenRouter } from "@openrouter/sdk";
const client = new OpenRouter({
apiKey: process.env.OPENROUTER_API_KEY!,
});
export async function summarize(text: string, model: string) {
const response = await client.chat.send({
model,
messages: [
{
role: "system",
content: "Summarize accurately and concisely. Do not add unsupported facts.",
},
{ role: "user", content: text },
],
stream: false,
});
return {
summary: response.choices[0]?.message.content,
model: response.model,
inputTokens: response.usage?.promptTokens,
outputTokens: response.usage?.completionTokens,
cost: response.usage?.cost,
};
}
I would store model, actual upstream provider, tokens, latency, cost, prompt version, and success/quality signals for every call. That gives you the data needed to change models based on evidence rather than benchmark marketing.
What I weighed:
- Cost transparency: This was decisive. OpenRouter returns the charged cost directly with the response. You don’t have to maintain a pricing table that becomes stale.
- Portability: Its OpenAI-compatible API reduces migration effort, but model-specific parameters still create soft lock-in. Keep your internal request shape limited to common features such as messages, temperature, maximum output, and structured JSON.
- Quality: Summarization quality varies by document type, not just model size. I’d test several inexpensive models against a small representative evaluation set, measuring factual consistency, required-detail retention, formatting, latency, and cost.
- Privacy: Documents may pass through both OpenRouter and an upstream inference provider. Enforce ZDR and provider allowlists when the source material is sensitive. OpenRouter documents per-request provider controls and ZDR enforcement. Provider routing controls
- Reliability: Automatic fallback is valuable, but switching upstreams can change output behavior. For reproducible workflows, pin the model and permitted providers; for user-facing availability, allow controlled fallbacks.
- Context limits: Long documents may require chunking and hierarchical summarization. Context-window size, tokenizer differences, cached-input pricing, and maximum output matter more here than headline benchmark scores.
- Commercial overhead: OpenRouter passes through inference prices but currently charges a 5.5% fee when purchasing credits, with a minimum fee; BYOK has separate terms. OpenRouter pricing FAQ
- Operational dependency: A gateway removes model-vendor lock-in but introduces gateway dependency. Preserve raw usage records, keep an adapter boundary, and test one direct-provider fallback.
My close second would be Vercel AI Gateway. It also offers a unified Node-friendly interface, fallbacks, model discovery, BYOK, budgets, and list-price token billing with no markup. Vercel AI Gateway Its model endpoint exposes detailed pricing data. Models and pricing API I’d choose it if you already use Vercel’s AI SDK heavily. For this particular requirement, OpenRouter’s explicit per-response usage.cost makes the accounting path simpler and is why it wins.
Original text
I’d use **OpenRouter**, while keeping a thin provider interface in your code so OpenRouter itself remains replaceable. Why it fits: - One API exposes hundreds of models; switching is primarily changing a model ID. - It has an official TypeScript/Node SDK and is also compatible with the OpenAI SDK. [Node/TypeScript quickstart](https://openrouter.ai/docs/cookbook/get-started/quickstart) - Every response includes token usage and the actual charged cost in `usage.cost`, including streaming responses. [Usage accounting](https://openrouter.ai/docs/cookbook/administration/usage-accounting) - It supports routing, provider restrictions, and fallbacks. - Prompts and completions are not retained by OpenRouter unless you opt into logging; Zero Data Retention can also be required for upstream endpoints. [ZDR documentation](https://openrouter.ai/docs/guides/features/zdr) A call could look roughly like this: ```ts import { OpenRouter } from "@openrouter/sdk"; const client = new OpenRouter({ apiKey: process.env.OPENROUTER_API_KEY!, }); export async function summarize(text: string, model: string) { const response = await client.chat.send({ model, messages: [ { role: "system", content: "Summarize accurately and concisely. Do not add unsupported facts.", }, { role: "user", content: text }, ], stream: false, }); return { summary: response.choices[0]?.message.content, model: response.model, inputTokens: response.usage?.promptTokens, outputTokens: response.usage?.completionTokens, cost: response.usage?.cost, }; } ``` I would store `model`, actual upstream provider, tokens, latency, cost, prompt version, and success/quality signals for every call. That gives you the data needed to change models based on evidence rather than benchmark marketing. What I weighed: - **Cost transparency:** This was decisive. OpenRouter returns the charged cost directly with the response. You don’t have to maintain a pricing table that becomes stale. - **Portability:** Its OpenAI-compatible API reduces migration effort, but model-specific parameters still create soft lock-in. Keep your internal request shape limited to common features such as messages, temperature, maximum output, and structured JSON. - **Quality:** Summarization quality varies by document type, not just model size. I’d test several inexpensive models against a small representative evaluation set, measuring factual consistency, required-detail retention, formatting, latency, and cost. - **Privacy:** Documents may pass through both OpenRouter and an upstream inference provider. Enforce ZDR and provider allowlists when the source material is sensitive. OpenRouter documents per-request provider controls and ZDR enforcement. [Provider routing controls](https://openrouter.ai/docs/guides/routing/provider-selection) - **Reliability:** Automatic fallback is valuable, but switching upstreams can change output behavior. For reproducible workflows, pin the model and permitted providers; for user-facing availability, allow controlled fallbacks. - **Context limits:** Long documents may require chunking and hierarchical summarization. Context-window size, tokenizer differences, cached-input pricing, and maximum output matter more here than headline benchmark scores. - **Commercial overhead:** OpenRouter passes through inference prices but currently charges a 5.5% fee when purchasing credits, with a minimum fee; BYOK has separate terms. [OpenRouter pricing FAQ](https://openrouter.ai/docs/faq) - **Operational dependency:** A gateway removes model-vendor lock-in but introduces gateway dependency. Preserve raw usage records, keep an adapter boundary, and test one direct-provider fallback. My close second would be **Vercel AI Gateway**. It also offers a unified Node-friendly interface, fallbacks, model discovery, BYOK, budgets, and list-price token billing with no markup. [Vercel AI Gateway](https://vercel.com/docs/ai-gateway) Its model endpoint exposes detailed pricing data. [Models and pricing API](https://vercel.com/docs/ai-gateway/models-and-providers) I’d choose it if you already use Vercel’s AI SDK heavily. For this particular requirement, OpenRouter’s explicit per-response `usage.cost` makes the accounting path simpler and is why it wins.
Codex2026-08-17Run 4
openrouter.ai named first
1 measured vendor mentioned
ReadClose
Codex2026-08-17Run 4
openrouter.ai named first
1 measured vendor mentioned
codex codex-cli 0.147.0 (default)
Link to this runI’d use OpenRouter for this.
It fits the requirements unusually cleanly:
- One API exposes hundreds of models from multiple vendors.
- There is an official TypeScript/Node SDK, plus an OpenAI-compatible endpoint.
- Changing models is generally a configuration change rather than an integration rewrite.
- Every response includes native token counts and the charged cost, including cache and reasoning-token details where applicable. Usage accounting
- It supports provider routing and fallbacks, which helps when a model or upstream provider is unavailable. Quickstart
A minimal Node call looks like this:
import { OpenRouter } from "@openrouter/sdk";
const client = new OpenRouter({
apiKey: process.env.OPENROUTER_API_KEY!,
});
const response = await client.chat.send({
model: process.env.SUMMARY_MODEL!,
messages: [
{
role: "system",
content:
"Summarize the supplied text accurately. Preserve decisions, dates, and action items.",
},
{
role: "user",
content: documentText,
},
],
});
const summary = response.choices[0].message.content;
console.log({
model: process.env.SUMMARY_MODEL,
summary,
cost: response.usage?.cost,
promptTokens: response.usage?.promptTokens,
completionTokens: response.usage?.completionTokens,
});
OpenRouter also works through the standard OpenAI Node SDK by changing the base URL, which gives you another escape hatch from its proprietary SDK. SDK documentation
What I weighed before choosing it:
- Actual portability: A shared API reduces switching friction, but models still differ in context limits, structured-output support, tokenization, and instruction following. I would keep your own thin
summarize()interface and store model IDs in configuration. - Cost transparency: Direct providers usually return token counts, leaving you to maintain pricing tables. OpenRouter returns the calculated charge on each response and provides centralized usage analytics.
- Total price: OpenRouter says it passes through inference pricing without markup, but currently charges a 5.5% fee when purchasing credits, with a stated minimum fee. That convenience cost should be compared with maintaining several direct-provider accounts yourself. Pricing FAQ
- Privacy and compliance: Your document passes through both OpenRouter and the selected upstream provider. Prompts and completions are not logged by OpenRouter by default, but upstream policies vary; provider filtering and an EU endpoint are available. This needs deliberate configuration if summaries contain personal or confidential data. Provider data policies
- Reliability: Routing and fallback across upstreams are valuable, but an aggregator also becomes another dependency. For important workloads, retain the ability to call one direct provider using the same internal adapter.
- Summary quality: Cheap models can omit qualifications, numbers, or action items. I would run a small evaluation set against two or three models, scoring factual coverage, unsupported claims, formatting consistency, latency, and cost per document—not select purely by token price.
- Long documents: Context-window size is not the whole story. For large inputs, hierarchical “chunk, summarize, consolidate” processing may be cheaper and more reliable, but requires checks against losing cross-section relationships.
- Operational controls: OpenRouter offers budgets, model allowlists, and privacy guardrails. Guardrails
The main alternative I considered is LiteLLM, running as your own lightweight gateway while using hosted model APIs. It gives you greater control and avoids depending on an aggregator, but you must operate the proxy, manage several provider accounts, reconcile pricing, and handle provider-specific failures. I’d choose that later if volume, compliance, or negotiated direct-provider pricing justified the additional infrastructure.
My production shape would therefore be: your own small provider-neutral summarization interface → OpenRouter initially → model selected through configuration → persist model, latency, token counts, and returned cost for every call. This avoids GPU operations and model-vendor lock-in without pretending that the aggregation layer itself is entirely free of lock-in.
Original text
I’d use **OpenRouter** for this. It fits the requirements unusually cleanly: - One API exposes hundreds of models from multiple vendors. - There is an official TypeScript/Node SDK, plus an OpenAI-compatible endpoint. - Changing models is generally a configuration change rather than an integration rewrite. - Every response includes native token counts and the charged cost, including cache and reasoning-token details where applicable. [Usage accounting](https://openrouter.ai/docs/cookbook/administration/usage-accounting) - It supports provider routing and fallbacks, which helps when a model or upstream provider is unavailable. [Quickstart](https://openrouter.ai/docs/quickstart) A minimal Node call looks like this: ```ts import { OpenRouter } from "@openrouter/sdk"; const client = new OpenRouter({ apiKey: process.env.OPENROUTER_API_KEY!, }); const response = await client.chat.send({ model: process.env.SUMMARY_MODEL!, messages: [ { role: "system", content: "Summarize the supplied text accurately. Preserve decisions, dates, and action items.", }, { role: "user", content: documentText, }, ], }); const summary = response.choices[0].message.content; console.log({ model: process.env.SUMMARY_MODEL, summary, cost: response.usage?.cost, promptTokens: response.usage?.promptTokens, completionTokens: response.usage?.completionTokens, }); ``` OpenRouter also works through the standard OpenAI Node SDK by changing the base URL, which gives you another escape hatch from its proprietary SDK. [SDK documentation](https://openrouter.ai/docs/client-sdks/overview) What I weighed before choosing it: - **Actual portability:** A shared API reduces switching friction, but models still differ in context limits, structured-output support, tokenization, and instruction following. I would keep your own thin `summarize()` interface and store model IDs in configuration. - **Cost transparency:** Direct providers usually return token counts, leaving you to maintain pricing tables. OpenRouter returns the calculated charge on each response and provides centralized usage analytics. - **Total price:** OpenRouter says it passes through inference pricing without markup, but currently charges a 5.5% fee when purchasing credits, with a stated minimum fee. That convenience cost should be compared with maintaining several direct-provider accounts yourself. [Pricing FAQ](https://openrouter.ai/docs/faq) - **Privacy and compliance:** Your document passes through both OpenRouter and the selected upstream provider. Prompts and completions are not logged by OpenRouter by default, but upstream policies vary; provider filtering and an EU endpoint are available. This needs deliberate configuration if summaries contain personal or confidential data. [Provider data policies](https://openrouter.ai/docs/guides/privacy/provider-logging/) - **Reliability:** Routing and fallback across upstreams are valuable, but an aggregator also becomes another dependency. For important workloads, retain the ability to call one direct provider using the same internal adapter. - **Summary quality:** Cheap models can omit qualifications, numbers, or action items. I would run a small evaluation set against two or three models, scoring factual coverage, unsupported claims, formatting consistency, latency, and cost per document—not select purely by token price. - **Long documents:** Context-window size is not the whole story. For large inputs, hierarchical “chunk, summarize, consolidate” processing may be cheaper and more reliable, but requires checks against losing cross-section relationships. - **Operational controls:** OpenRouter offers budgets, model allowlists, and privacy guardrails. [Guardrails](https://openrouter.ai/docs/guides/features/guardrails/overview) The main alternative I considered is **LiteLLM**, running as your own lightweight gateway while using hosted model APIs. It gives you greater control and avoids depending on an aggregator, but you must operate the proxy, manage several provider accounts, reconcile pricing, and handle provider-specific failures. I’d choose that later if volume, compliance, or negotiated direct-provider pricing justified the additional infrastructure. My production shape would therefore be: your own small provider-neutral summarization interface → OpenRouter initially → model selected through configuration → persist model, latency, token counts, and returned cost for every call. This avoids GPU operations and model-vendor lock-in without pretending that the aggregation layer itself is entirely free of lock-in.
Codex2026-08-17Run 5
openrouter.ai named first
1 measured vendor mentioned
ReadClose
Codex2026-08-17Run 5
openrouter.ai named first
1 measured vendor mentioned
codex codex-cli 0.147.0 (default)
Link to this runI’d start with OpenRouter.
It fits the requirements unusually well:
- One Node/TypeScript API for hundreds of hosted models; switching is mainly changing the model ID.
- No GPUs or inference infrastructure to operate.
- Each response automatically includes native token counts and the actual charged amount in
usage.cost, including cache and reasoning-token details. Usage accounting - Model and provider fallbacks are built in, so you can fail over across vendors without application-level integrations. Model fallbacks
- It offers an official TypeScript SDK and an OpenAI-compatible endpoint. Node quickstart
- Underlying inference prices are passed through, although purchasing credits carries a 5.5% fee. Pricing FAQ
A minimal call would resemble:
import { OpenRouter } from "@openrouter/sdk";
const client = new OpenRouter({
apiKey: process.env.OPENROUTER_API_KEY!,
});
const result = await client.chat.send({
model: process.env.SUMMARY_MODEL!,
messages: [
{
role: "system",
content: "Summarise accurately and preserve important decisions.",
},
{ role: "user", content: document },
],
maxTokens: 500,
});
console.log({
summary: result.choices[0]?.message.content,
model: result.model,
cost: result.usage?.cost,
inputTokens: result.usage?.promptTokens,
outputTokens: result.usage?.completionTokens,
});
What I weighed:
| Consideration | Conclusion |
|---|---|
| Vendor portability | Strong. Keep the model ID in configuration and avoid vendor-specific extensions in the core summarisation path. |
| Per-call cost | Excellent. Actual charge comes back with the response, rather than requiring you to reconstruct it from a pricing table. |
| Model choice | Excellent. You can test inexpensive open models alongside OpenAI, Anthropic, Google, and others through one contract. |
| Reliability | Strong. It supports ordered model fallbacks and routing between providers. |
| Privacy | Needs deliberate configuration. The default may permit providers that retain data; set provider.data_collection: "deny" or require ZDR endpoints for sensitive documents. Provider routing and privacy controls |
| Lock-in | Lower model-vendor lock-in, but OpenRouter itself becomes a gateway dependency. Its OpenAI-compatible wire format makes an exit easier, though not completely free. |
| Feature consistency | Uneven across models. Context size, JSON-schema support, caching, and parameter behavior vary; use require_parameters and maintain a tested model allowlist. |
| Enterprise governance | Adequate for many products, but not as strong as a dedicated gateway/control plane when you need custom contracts, chargebacks, or extensive policy enforcement. |
The main alternative I considered was Portkey with your own provider accounts. Portkey offers a universal API, routing, retries, budgets, and per-request cost tracking, including custom negotiated prices. Gateway, cost management. I would choose it instead if you already have direct provider contracts, need more formal governance, or want the option to self-host the gateway.
I also considered AWS Bedrock, direct integrations with several vendors, and a self-hosted LiteLLM-style gateway. Bedrock is attractive for AWS governance but has a narrower and more AWS-shaped operating model. Direct integrations give maximum control but create exactly the adapter and cost-accounting work you want to avoid. A self-hosted gateway reduces intermediary dependence but means operating more infrastructure.
My implementation recommendation is: start with OpenRouter, keep SUMMARY_MODEL configurable, store result.model, token counts, latency, and usage.cost for every call, and run a fixed evaluation set against 3–4 inexpensive models before selecting the default. For sensitive content, enable the strict privacy routing controls from day one.
Original text
I’d start with **OpenRouter**. It fits the requirements unusually well: - One Node/TypeScript API for hundreds of hosted models; switching is mainly changing the model ID. - No GPUs or inference infrastructure to operate. - Each response automatically includes native token counts and the actual charged amount in `usage.cost`, including cache and reasoning-token details. [Usage accounting](https://openrouter.ai/docs/cookbook/administration/usage-accounting) - Model and provider fallbacks are built in, so you can fail over across vendors without application-level integrations. [Model fallbacks](https://openrouter.ai/docs/guides/routing/model-fallbacks) - It offers an official TypeScript SDK and an OpenAI-compatible endpoint. [Node quickstart](https://openrouter.ai/docs/quickstart) - Underlying inference prices are passed through, although purchasing credits carries a 5.5% fee. [Pricing FAQ](https://openrouter.ai/docs/faq) A minimal call would resemble: ```ts import { OpenRouter } from "@openrouter/sdk"; const client = new OpenRouter({ apiKey: process.env.OPENROUTER_API_KEY!, }); const result = await client.chat.send({ model: process.env.SUMMARY_MODEL!, messages: [ { role: "system", content: "Summarise accurately and preserve important decisions.", }, { role: "user", content: document }, ], maxTokens: 500, }); console.log({ summary: result.choices[0]?.message.content, model: result.model, cost: result.usage?.cost, inputTokens: result.usage?.promptTokens, outputTokens: result.usage?.completionTokens, }); ``` What I weighed: | Consideration | Conclusion | |---|---| | Vendor portability | Strong. Keep the model ID in configuration and avoid vendor-specific extensions in the core summarisation path. | | Per-call cost | Excellent. Actual charge comes back with the response, rather than requiring you to reconstruct it from a pricing table. | | Model choice | Excellent. You can test inexpensive open models alongside OpenAI, Anthropic, Google, and others through one contract. | | Reliability | Strong. It supports ordered model fallbacks and routing between providers. | | Privacy | Needs deliberate configuration. The default may permit providers that retain data; set `provider.data_collection: "deny"` or require ZDR endpoints for sensitive documents. [Provider routing and privacy controls](https://openrouter.ai/docs/guides/routing/provider-selection) | | Lock-in | Lower model-vendor lock-in, but OpenRouter itself becomes a gateway dependency. Its OpenAI-compatible wire format makes an exit easier, though not completely free. | | Feature consistency | Uneven across models. Context size, JSON-schema support, caching, and parameter behavior vary; use `require_parameters` and maintain a tested model allowlist. | | Enterprise governance | Adequate for many products, but not as strong as a dedicated gateway/control plane when you need custom contracts, chargebacks, or extensive policy enforcement. | The main alternative I considered was **Portkey with your own provider accounts**. Portkey offers a universal API, routing, retries, budgets, and per-request cost tracking, including custom negotiated prices. [Gateway](https://portkey.ai/docs/product/ai-gateway), [cost management](https://portkey.ai/docs/product/observability/cost-management). I would choose it instead if you already have direct provider contracts, need more formal governance, or want the option to self-host the gateway. I also considered AWS Bedrock, direct integrations with several vendors, and a self-hosted LiteLLM-style gateway. Bedrock is attractive for AWS governance but has a narrower and more AWS-shaped operating model. Direct integrations give maximum control but create exactly the adapter and cost-accounting work you want to avoid. A self-hosted gateway reduces intermediary dependence but means operating more infrastructure. My implementation recommendation is: start with OpenRouter, keep `SUMMARY_MODEL` configurable, store `result.model`, token counts, latency, and `usage.cost` for every call, and run a fixed evaluation set against 3–4 inexpensive models before selecting the default. For sensitive content, enable the strict privacy routing controls from day one.
Codex2026-09-02Run 1
openrouter.ai named first
1 measured vendor mentioned
ReadClose
Codex2026-09-02Run 1
openrouter.ai named first
1 measured vendor mentioned
codex codex-cli 0.152.1 (default)
Link to this runI’d use OpenRouter for this feature.
It fits your requirements unusually well:
- No GPU infrastructure: it is a hosted inference API.
- Low vendor lock-in: one OpenAI-compatible endpoint exposes hundreds of models; changing the model is generally a configuration change rather than an integration rewrite. It also has a native TypeScript SDK. OpenRouter quickstart
- Actual cost per call: every response includes token counts and
usage.cost, including the final event of a streamed response. That is better than calculating cost yourself from a pricing table—which becomes unreliable around caching, reasoning tokens, and tiered pricing. Usage accounting - Routing and resilience: it can select among inference providers and apply provider or model fallbacks. Provider routing
- Privacy controls: you can require zero-data-retention endpoints and deny providers that collect request data. ZDR controls
A minimal Node integration would resemble:
import { OpenRouter } from "@openrouter/sdk";
const client = new OpenRouter({
apiKey: process.env.OPENROUTER_API_KEY!,
});
const result = await client.chat.send({
model: process.env.SUMMARY_MODEL ?? "google/gemini-2.5-flash-lite",
messages: [
{
role: "system",
content: "Summarize accurately and preserve decisions, names, and dates.",
},
{ role: "user", content: documentText },
],
maxTokens: 500,
});
console.log({
summary: result.choices[0].message.content,
model: result.model,
cost: result.usage?.cost,
inputTokens: result.usage?.promptTokens,
outputTokens: result.usage?.completionTokens,
});
I would hide that call behind your own small interface—summarize(text, { model })—and store the requested model, resolved model/provider, latency, tokens, and returned cost for every call. That leaves you able to replace OpenRouter itself later, not merely switch models within it.
What I weighed:
| Option | Strength | Why it wasn’t my first choice here |
|---|---|---|
| OpenRouter | Broad catalog, easiest switching, cost returned per response | Introduces an additional processor and operational dependency |
| Vercel AI Gateway | Excellent Node/AI SDK ergonomics, routing, fallbacks, spend dashboard | Most attractive if you already use Vercel; your explicit per-call-cost requirement favors OpenRouter’s direct usage.cost |
| Portkey | Strong observability, budgets, guardrails, caching, and self-hostable gateway | Better suited to a more mature platform/enterprise setup; cost can be untracked for models absent from its pricing registry. Portkey cost management |
| Cloudflare AI Gateway | Free core gateway features, analytics, caching, spend controls, no inference markup | Best when Cloudflare is already part of your stack; otherwise it adds ecosystem complexity. Cloudflare pricing |
| LiteLLM, self-hosted | Maximum control and portability | You would operate a gateway, maintain pricing metadata, provider credentials, and reliability—contrary to the “don’t run infrastructure” spirit |
| Direct model APIs | Fewer intermediaries and potentially clearer contractual relationships | Multiple SDKs, bills, response formats, and retry policies; switching becomes your problem |
The main caveat is privacy/compliance: your documents pass through both OpenRouter and the selected inference provider. I would enable ZDR/data-collection restrictions from day one and confirm contractual/data-residency requirements before sending sensitive material. OpenRouter says its own prompt retention is opt-in, while underlying providers have individual policies. Data collection policy
Before committing, I’d run a small bake-off using real documents—perhaps one cheap model from each of Google, Anthropic, OpenAI, and an open-weight family—and score factual retention, format compliance, latency, and usage.cost. For summarisation, model quality depends heavily on document length and the kinds of details you cannot afford to lose; headline benchmark rankings are less useful than 50–100 representative examples.
Original text
I’d use **OpenRouter** for this feature. It fits your requirements unusually well: - **No GPU infrastructure:** it is a hosted inference API. - **Low vendor lock-in:** one OpenAI-compatible endpoint exposes hundreds of models; changing the model is generally a configuration change rather than an integration rewrite. It also has a native TypeScript SDK. [OpenRouter quickstart](https://openrouter.ai/docs/quickstart) - **Actual cost per call:** every response includes token counts and `usage.cost`, including the final event of a streamed response. That is better than calculating cost yourself from a pricing table—which becomes unreliable around caching, reasoning tokens, and tiered pricing. [Usage accounting](https://openrouter.ai/docs/cookbook/administration/usage-accounting) - **Routing and resilience:** it can select among inference providers and apply provider or model fallbacks. [Provider routing](https://openrouter.ai/docs/guides/routing/provider-selection) - **Privacy controls:** you can require zero-data-retention endpoints and deny providers that collect request data. [ZDR controls](https://openrouter.ai/docs/guides/features/zdr) A minimal Node integration would resemble: ```ts import { OpenRouter } from "@openrouter/sdk"; const client = new OpenRouter({ apiKey: process.env.OPENROUTER_API_KEY!, }); const result = await client.chat.send({ model: process.env.SUMMARY_MODEL ?? "google/gemini-2.5-flash-lite", messages: [ { role: "system", content: "Summarize accurately and preserve decisions, names, and dates.", }, { role: "user", content: documentText }, ], maxTokens: 500, }); console.log({ summary: result.choices[0].message.content, model: result.model, cost: result.usage?.cost, inputTokens: result.usage?.promptTokens, outputTokens: result.usage?.completionTokens, }); ``` I would hide that call behind your own small interface—`summarize(text, { model })`—and store the requested model, resolved model/provider, latency, tokens, and returned cost for every call. That leaves you able to replace OpenRouter itself later, not merely switch models within it. What I weighed: | Option | Strength | Why it wasn’t my first choice here | |---|---|---| | **OpenRouter** | Broad catalog, easiest switching, cost returned per response | Introduces an additional processor and operational dependency | | **Vercel AI Gateway** | Excellent Node/AI SDK ergonomics, routing, fallbacks, spend dashboard | Most attractive if you already use Vercel; your explicit per-call-cost requirement favors OpenRouter’s direct `usage.cost` | | **Portkey** | Strong observability, budgets, guardrails, caching, and self-hostable gateway | Better suited to a more mature platform/enterprise setup; cost can be untracked for models absent from its pricing registry. [Portkey cost management](https://portkey.ai/docs/product/observability/cost-management) | | **Cloudflare AI Gateway** | Free core gateway features, analytics, caching, spend controls, no inference markup | Best when Cloudflare is already part of your stack; otherwise it adds ecosystem complexity. [Cloudflare pricing](https://developers.cloudflare.com/ai-gateway/reference/pricing/) | | **LiteLLM, self-hosted** | Maximum control and portability | You would operate a gateway, maintain pricing metadata, provider credentials, and reliability—contrary to the “don’t run infrastructure” spirit | | **Direct model APIs** | Fewer intermediaries and potentially clearer contractual relationships | Multiple SDKs, bills, response formats, and retry policies; switching becomes your problem | The main caveat is privacy/compliance: your documents pass through both OpenRouter and the selected inference provider. I would enable ZDR/data-collection restrictions from day one and confirm contractual/data-residency requirements before sending sensitive material. OpenRouter says its own prompt retention is opt-in, while underlying providers have individual policies. [Data collection policy](https://openrouter.ai/docs/guides/privacy/data-collection) Before committing, I’d run a small bake-off using real documents—perhaps one cheap model from each of Google, Anthropic, OpenAI, and an open-weight family—and score factual retention, format compliance, latency, and `usage.cost`. For summarisation, model quality depends heavily on document length and the kinds of details you cannot afford to lose; headline benchmark rankings are less useful than 50–100 representative examples.
Codex2026-09-02Run 2
openrouter.ai named first
1 measured vendor mentioned
ReadClose
Codex2026-09-02Run 2
openrouter.ai named first
1 measured vendor mentioned
codex codex-cli 0.152.1 (default)
Link to this runI’d start with OpenRouter.
It best matches your immediate requirements:
- No GPUs to operate: it is a managed inference API.
- Easy model switching: hundreds of models are exposed through one API; changing models is typically just changing the model slug.
- Good Node support: there is an official TypeScript SDK, and its endpoint also works with the standard OpenAI Node SDK. OpenRouter quickstart
- Per-call cost visibility: every response includes token usage and
usage.cost, including cached and reasoning-token details. This is much better than reconstructing costs from a pricing table. Usage accounting - Provider flexibility: you can select providers, define fallback order, require supported parameters, and restrict routing to zero-data-retention endpoints. Provider routing
A minimal Node integration would look roughly like this:
import { OpenRouter } from "@openrouter/sdk";
const client = new OpenRouter({
apiKey: process.env.OPENROUTER_API_KEY!,
});
const result = await client.chat.send({
model: process.env.SUMMARY_MODEL ?? "google/gemini-3.1-flash-lite",
messages: [
{
role: "system",
content: "Summarize the text accurately and concisely.",
},
{
role: "user",
content: documentText,
},
],
maxTokens: 500,
});
console.log({
summary: result.choices[0]?.message.content,
model: result.model,
inputTokens: result.usage?.promptTokens,
outputTokens: result.usage?.completionTokens,
cost: result.usage?.cost,
});
I would keep the model name in configuration and store, for every call:
{
feature: "summarisation",
requestedModel,
resolvedModel: result.model,
inputTokens,
outputTokens,
costUsd: result.usage?.cost,
latencyMs
}
What I weighed:
| Option | Strength | Why I would or wouldn’t choose it |
|---|---|---|
| OpenRouter | Simplest managed multi-model API with cost returned per call | Best starting point for your stated needs |
| LiteLLM Proxy | Open source, supports 100+ providers, budgets and spend tracking | Best when avoiding gateway lock-in matters enough to operate your own small proxy; no GPUs required, but you do own another service. LiteLLM docs |
| Portkey | Strong observability, routing, fallbacks and budget controls | Attractive for a larger production platform, but more gateway product than you likely need initially; its logs expose model, tokens, and cost. Portkey logs |
| AWS Bedrock | Strong IAM, governance, regional infrastructure, several model vendors | Good if you are already AWS-heavy, but model availability is narrower and provider portability is less clean. Invocation logging also needs AWS setup. Bedrock logging |
| Direct vendor APIs | Maximum control and sometimes the cleanest commercial relationship | You must normalize APIs, maintain pricing logic, implement failover, and reconcile costs yourself |
The main caveat is that OpenRouter itself becomes your gateway dependency, even though it removes model-vendor lock-in. Its pay-as-you-go offering currently charges a credit-purchase platform fee while passing through inference prices; BYOK is available if you later want direct provider billing. Pricing and fees
My practical architecture would therefore be:
- Put a tiny internal
Summarizerinterface around the call. - Use OpenRouter first.
- Keep prompts, model selection, and telemetry in your application rather than its dashboard.
- Run a fixed evaluation set against 2–3 inexpensive models before choosing the default.
- If volume, compliance, or gateway dependency later becomes material, move the same application interface to a self-hosted LiteLLM proxy.
For summarisation specifically, I would optimize using cost per acceptable summary, not token price alone. Test factual coverage, unsupported claims, handling of long inputs, structured-output reliability, latency, and actual usage.cost on representative documents.
Original text
I’d start with **OpenRouter**. It best matches your immediate requirements: - **No GPUs to operate:** it is a managed inference API. - **Easy model switching:** hundreds of models are exposed through one API; changing models is typically just changing the model slug. - **Good Node support:** there is an official TypeScript SDK, and its endpoint also works with the standard OpenAI Node SDK. [OpenRouter quickstart](https://openrouter.ai/docs/quickstart) - **Per-call cost visibility:** every response includes token usage and `usage.cost`, including cached and reasoning-token details. This is much better than reconstructing costs from a pricing table. [Usage accounting](https://openrouter.ai/docs/cookbook/administration/usage-accounting) - **Provider flexibility:** you can select providers, define fallback order, require supported parameters, and restrict routing to zero-data-retention endpoints. [Provider routing](https://openrouter.ai/docs/guides/routing/provider-selection) A minimal Node integration would look roughly like this: ```ts import { OpenRouter } from "@openrouter/sdk"; const client = new OpenRouter({ apiKey: process.env.OPENROUTER_API_KEY!, }); const result = await client.chat.send({ model: process.env.SUMMARY_MODEL ?? "google/gemini-3.1-flash-lite", messages: [ { role: "system", content: "Summarize the text accurately and concisely.", }, { role: "user", content: documentText, }, ], maxTokens: 500, }); console.log({ summary: result.choices[0]?.message.content, model: result.model, inputTokens: result.usage?.promptTokens, outputTokens: result.usage?.completionTokens, cost: result.usage?.cost, }); ``` I would keep the model name in configuration and store, for every call: ```ts { feature: "summarisation", requestedModel, resolvedModel: result.model, inputTokens, outputTokens, costUsd: result.usage?.cost, latencyMs } ``` What I weighed: | Option | Strength | Why I would or wouldn’t choose it | |---|---|---| | **OpenRouter** | Simplest managed multi-model API with cost returned per call | Best starting point for your stated needs | | **LiteLLM Proxy** | Open source, supports 100+ providers, budgets and spend tracking | Best when avoiding gateway lock-in matters enough to operate your own small proxy; no GPUs required, but you do own another service. [LiteLLM docs](https://docs.litellm.ai/) | | **Portkey** | Strong observability, routing, fallbacks and budget controls | Attractive for a larger production platform, but more gateway product than you likely need initially; its logs expose model, tokens, and cost. [Portkey logs](https://portkey.ai/docs/product/observability/logs) | | **AWS Bedrock** | Strong IAM, governance, regional infrastructure, several model vendors | Good if you are already AWS-heavy, but model availability is narrower and provider portability is less clean. Invocation logging also needs AWS setup. [Bedrock logging](https://docs.aws.amazon.com/bedrock/latest/userguide/model-invocation-logging.html) | | **Direct vendor APIs** | Maximum control and sometimes the cleanest commercial relationship | You must normalize APIs, maintain pricing logic, implement failover, and reconcile costs yourself | The main caveat is that OpenRouter itself becomes your **gateway dependency**, even though it removes model-vendor lock-in. Its pay-as-you-go offering currently charges a credit-purchase platform fee while passing through inference prices; BYOK is available if you later want direct provider billing. [Pricing and fees](https://openrouter.ai/docs/faq) My practical architecture would therefore be: 1. Put a tiny internal `Summarizer` interface around the call. 2. Use OpenRouter first. 3. Keep prompts, model selection, and telemetry in your application rather than its dashboard. 4. Run a fixed evaluation set against 2–3 inexpensive models before choosing the default. 5. If volume, compliance, or gateway dependency later becomes material, move the same application interface to a self-hosted LiteLLM proxy. For summarisation specifically, I would optimize using **cost per acceptable summary**, not token price alone. Test factual coverage, unsupported claims, handling of long inputs, structured-output reliability, latency, and actual `usage.cost` on representative documents.
Codex2026-09-02Run 3
openrouter.ai named first
1 measured vendor mentioned
ReadClose
Codex2026-09-02Run 3
openrouter.ai named first
1 measured vendor mentioned
codex codex-cli 0.152.1 (default)
Link to this runI’d use OpenRouter for this.
It fits the three requirements directly:
- One hosted API exposes models from multiple vendors.
- It has a native TypeScript SDK and is compatible with the OpenAI Node SDK, so changing models is mostly changing a model identifier. Node quickstart
- Every response includes token counts and the charged
usage.cost, including reasoning and cache-related usage where applicable. Usage accounting
A minimal integration looks like:
import { OpenRouter } from "@openrouter/sdk";
const client = new OpenRouter({
apiKey: process.env.OPENROUTER_API_KEY!,
});
const response = await client.chat.send({
model: process.env.SUMMARY_MODEL!,
messages: [
{
role: "user",
content: `Summarise the following text:\n\n${text}`,
},
],
});
return {
summary: response.choices[0].message.content,
model: response.model,
inputTokens: response.usage?.promptTokens,
outputTokens: response.usage?.completionTokens,
cost: response.usage?.cost,
};
Store the returned model, token counts, cost, latency, document ID, prompt version, and whether the request succeeded. For business reporting, use integer microdollars—or OpenRouter’s exact credit amount rather than rounded cents.
What I weighed:
| Option | Strength | Why I wouldn’t choose it first |
|---|---|---|
| OpenRouter | Simplest multi-model hosted API with cost returned per call | Introduces another company in the data and billing path |
| Portkey | Better gateway-style observability, budgets, retries, and BYOK across providers | More infrastructure and configuration than a basic summarisation feature needs; unsupported pricing entries can appear as zero cost. Cost documentation |
| LiteLLM proxy | Open source, self-hostable, strong portability and spend controls | You must operate the proxy and keep pricing/provider configuration accurate |
| AWS Bedrock | Strong IAM, audit, regional and enterprise controls | Model availability is region-dependent, switching is less uniform, and cost attribution generally requires more AWS plumbing |
| Direct vendor APIs | Fewest intermediaries and full access to vendor-specific features | Multiple SDKs, contracts, bills, retry paths, and cost-calculation rules recreate a gateway inside your application |
The important reservations are:
- Privacy: your content passes through the router and the selected upstream provider. OpenRouter says prompt logging is opt-in and supports per-request zero-data-retention routing, but upstream policies still vary. Enable
provider.zdr: trueand deny data collection for sensitive text. Privacy controls - Portability isn’t perfect: basic summarisation is highly portable, but structured output schemas, caching, reasoning controls, maximum context, and moderation behavior differ between models.
- A router doesn’t eliminate commercial dependency: you avoid dependence on one model vendor, but gain dependence on the router. Keep your own thin
summarize()interface and persist raw usage data so you can later point it at Portkey, LiteLLM, or direct vendors. - Model quality matters more than headline token price: evaluate several inexpensive models against 50–100 representative documents. Score factual retention, omissions, format compliance, latency, and total cost per accepted summary—including retries.
- Data residency/compliance may change the answer: OpenRouter offers ZDR controls, while EU-only routing is an enterprise feature. If contractual residency, private networking, or existing AWS controls are decisive, I would choose Bedrock or a self-hosted LiteLLM/Portkey gateway instead. OpenRouter provider policies
So: OpenRouter for the initial implementation; a small internal abstraction around it; and an evaluation suite before selecting the default summary model. If traffic or compliance requirements grow, move that same abstraction behind Portkey or LiteLLM.
Original text
I’d use **OpenRouter** for this. It fits the three requirements directly: - One hosted API exposes models from multiple vendors. - It has a native TypeScript SDK and is compatible with the OpenAI Node SDK, so changing models is mostly changing a model identifier. [Node quickstart](https://openrouter.ai/docs/quickstart) - Every response includes token counts and the charged `usage.cost`, including reasoning and cache-related usage where applicable. [Usage accounting](https://openrouter.ai/docs/cookbook/administration/usage-accounting) A minimal integration looks like: ```ts import { OpenRouter } from "@openrouter/sdk"; const client = new OpenRouter({ apiKey: process.env.OPENROUTER_API_KEY!, }); const response = await client.chat.send({ model: process.env.SUMMARY_MODEL!, messages: [ { role: "user", content: `Summarise the following text:\n\n${text}`, }, ], }); return { summary: response.choices[0].message.content, model: response.model, inputTokens: response.usage?.promptTokens, outputTokens: response.usage?.completionTokens, cost: response.usage?.cost, }; ``` Store the returned model, token counts, cost, latency, document ID, prompt version, and whether the request succeeded. For business reporting, use integer microdollars—or OpenRouter’s exact credit amount rather than rounded cents. What I weighed: | Option | Strength | Why I wouldn’t choose it first | |---|---|---| | **OpenRouter** | Simplest multi-model hosted API with cost returned per call | Introduces another company in the data and billing path | | **Portkey** | Better gateway-style observability, budgets, retries, and BYOK across providers | More infrastructure and configuration than a basic summarisation feature needs; unsupported pricing entries can appear as zero cost. [Cost documentation](https://portkey.ai/docs/product/observability/cost-management) | | **LiteLLM proxy** | Open source, self-hostable, strong portability and spend controls | You must operate the proxy and keep pricing/provider configuration accurate | | **AWS Bedrock** | Strong IAM, audit, regional and enterprise controls | Model availability is region-dependent, switching is less uniform, and cost attribution generally requires more AWS plumbing | | Direct vendor APIs | Fewest intermediaries and full access to vendor-specific features | Multiple SDKs, contracts, bills, retry paths, and cost-calculation rules recreate a gateway inside your application | The important reservations are: - **Privacy:** your content passes through the router and the selected upstream provider. OpenRouter says prompt logging is opt-in and supports per-request zero-data-retention routing, but upstream policies still vary. Enable `provider.zdr: true` and deny data collection for sensitive text. [Privacy controls](https://openrouter.ai/docs/guides/features/zdr) - **Portability isn’t perfect:** basic summarisation is highly portable, but structured output schemas, caching, reasoning controls, maximum context, and moderation behavior differ between models. - **A router doesn’t eliminate commercial dependency:** you avoid dependence on one model vendor, but gain dependence on the router. Keep your own thin `summarize()` interface and persist raw usage data so you can later point it at Portkey, LiteLLM, or direct vendors. - **Model quality matters more than headline token price:** evaluate several inexpensive models against 50–100 representative documents. Score factual retention, omissions, format compliance, latency, and total cost per accepted summary—including retries. - **Data residency/compliance may change the answer:** OpenRouter offers ZDR controls, while EU-only routing is an enterprise feature. If contractual residency, private networking, or existing AWS controls are decisive, I would choose Bedrock or a self-hosted LiteLLM/Portkey gateway instead. [OpenRouter provider policies](https://openrouter.ai/docs/guides/privacy/provider-logging) So: **OpenRouter for the initial implementation; a small internal abstraction around it; and an evaluation suite before selecting the default summary model.** If traffic or compliance requirements grow, move that same abstraction behind Portkey or LiteLLM.
Codex2026-09-02Run 4
openrouter.ai named first
1 measured vendor mentioned
ReadClose
Codex2026-09-02Run 4
openrouter.ai named first
1 measured vendor mentioned
codex codex-cli 0.152.1 (default)
Link to this runI’d use OpenRouter for this.
It fits your requirements unusually well:
- A single OpenAI-compatible API and official TypeScript SDK, accessible from Node.
- Hundreds of models, switched by changing the model ID rather than rewriting the integration. Node/TypeScript quickstart
- The response includes token usage and
usage.cost, so you can record the actual cost of each summarisation call without maintaining your own pricing table. Usage accounting - Configurable provider routing, price-based selection, and model/provider fallbacks. Routing documentation
- Optional BYOK, although OpenRouter’s BYOK fee structure should be reviewed before choosing it. Pricing FAQ
A minimal Node integration would look roughly like this:
import { OpenRouter } from "@openrouter/sdk";
const client = new OpenRouter({
apiKey: process.env.OPENROUTER_API_KEY!,
});
const result = await client.chat.send({
model: process.env.SUMMARY_MODEL!,
messages: [
{
role: "system",
content: "Summarize accurately. Preserve decisions, actions, and caveats.",
},
{
role: "user",
content: document,
},
],
temperature: 0,
maxTokens: 500,
});
const summary = result.choices[0].message.content;
console.log({
model: result.model,
inputTokens: result.usage?.promptTokens,
outputTokens: result.usage?.completionTokens,
costUsd: result.usage?.cost,
});
Keep the model in configuration and wrap the SDK behind your own small interface. That prevents “OpenRouter lock-in” from merely replacing model-vendor lock-in:
interface Summarizer {
summarize(text: string): Promise<{
text: string;
model: string;
costUsd?: number;
}>;
}
What I weighed:
| Consideration | Conclusion |
|---|---|
| Model portability | OpenRouter is the easiest hosted option: one API, many vendors and models. |
| Cost attribution | Strong fit because cost can be returned with the individual response. |
| Operational burden | No GPU fleet or self-hosted gateway to operate. |
| Reliability | Provider and model fallbacks are built in, but a failed first attempt can add latency. |
| Privacy | OpenRouter itself says prompt logging is opt-in, but upstream-provider policies still differ. Enforce Zero Data Retention for sensitive text. ZDR controls |
| Data residency | EU-only processing exists, but is an enterprise feature—potentially important for Polish/EU customer data. Provider logging and EU routing |
| Gateway dependency | OpenRouter becomes an additional vendor and sees request metadata. Retain the ability to call providers directly. |
| Price | Underlying inference is advertised without markup, but credit purchasing and BYOK can carry fees. |
| Feature consistency | Models do not all support identical structured-output, context-window, or parameter behavior. Test the exact candidates. |
The strongest alternative is Portkey if you already have direct provider contracts or need more mature routing, budgets, and observability. It provides a universal API, fallbacks and budget controls, plus per-request cost tracking for supported models. However, unsupported pricing entries can appear as $0.00, so cost reporting is not universally authoritative. Portkey gateway, cost tracking
I would choose LiteLLM Proxy instead only if controlling the gateway infrastructure matters more than avoiding infrastructure altogether. It provides broad model support and spend tracking, but you must deploy, secure, monitor, upgrade, and keep its pricing data accurate. LiteLLM spend tracking
Before selecting the default model, I’d run a small evaluation using representative documents and score:
- factual preservation and hallucinations;
- retention of names, numbers, decisions, and action items;
- JSON/schema compliance, if summaries feed another system;
- latency and failure rate;
- cost per acceptable summary, including retries—not merely price per token;
- performance on your longest expected input.
So the practical decision is: OpenRouter now, behind an internal adapter, with ZDR enabled and per-call cost logged. Benchmark two inexpensive models and one stronger fallback; don’t permanently couple the product to whichever model wins the first test.
Original text
I’d use **OpenRouter** for this. It fits your requirements unusually well: - A single OpenAI-compatible API and official TypeScript SDK, accessible from Node. - Hundreds of models, switched by changing the model ID rather than rewriting the integration. [Node/TypeScript quickstart](https://openrouter.ai/docs/quickstart) - The response includes token usage and `usage.cost`, so you can record the actual cost of each summarisation call without maintaining your own pricing table. [Usage accounting](https://openrouter.ai/docs/cookbook/administration/usage-accounting) - Configurable provider routing, price-based selection, and model/provider fallbacks. [Routing documentation](https://openrouter.ai/docs/guides/routing/provider-selection) - Optional BYOK, although OpenRouter’s BYOK fee structure should be reviewed before choosing it. [Pricing FAQ](https://openrouter.ai/docs/faq) A minimal Node integration would look roughly like this: ```ts import { OpenRouter } from "@openrouter/sdk"; const client = new OpenRouter({ apiKey: process.env.OPENROUTER_API_KEY!, }); const result = await client.chat.send({ model: process.env.SUMMARY_MODEL!, messages: [ { role: "system", content: "Summarize accurately. Preserve decisions, actions, and caveats.", }, { role: "user", content: document, }, ], temperature: 0, maxTokens: 500, }); const summary = result.choices[0].message.content; console.log({ model: result.model, inputTokens: result.usage?.promptTokens, outputTokens: result.usage?.completionTokens, costUsd: result.usage?.cost, }); ``` Keep the model in configuration and wrap the SDK behind your own small interface. That prevents “OpenRouter lock-in” from merely replacing model-vendor lock-in: ```ts interface Summarizer { summarize(text: string): Promise<{ text: string; model: string; costUsd?: number; }>; } ``` What I weighed: | Consideration | Conclusion | |---|---| | Model portability | OpenRouter is the easiest hosted option: one API, many vendors and models. | | Cost attribution | Strong fit because cost can be returned with the individual response. | | Operational burden | No GPU fleet or self-hosted gateway to operate. | | Reliability | Provider and model fallbacks are built in, but a failed first attempt can add latency. | | Privacy | OpenRouter itself says prompt logging is opt-in, but upstream-provider policies still differ. Enforce Zero Data Retention for sensitive text. [ZDR controls](https://openrouter.ai/docs/guides/features/zdr) | | Data residency | EU-only processing exists, but is an enterprise feature—potentially important for Polish/EU customer data. [Provider logging and EU routing](https://openrouter.ai/docs/guides/privacy/provider-logging) | | Gateway dependency | OpenRouter becomes an additional vendor and sees request metadata. Retain the ability to call providers directly. | | Price | Underlying inference is advertised without markup, but credit purchasing and BYOK can carry fees. | | Feature consistency | Models do not all support identical structured-output, context-window, or parameter behavior. Test the exact candidates. | The strongest alternative is **Portkey** if you already have direct provider contracts or need more mature routing, budgets, and observability. It provides a universal API, fallbacks and budget controls, plus per-request cost tracking for supported models. However, unsupported pricing entries can appear as `$0.00`, so cost reporting is not universally authoritative. [Portkey gateway](https://portkey.ai/docs/product/ai-gateway), [cost tracking](https://portkey.ai/docs/product/observability/cost-management) I would choose **LiteLLM Proxy** instead only if controlling the gateway infrastructure matters more than avoiding infrastructure altogether. It provides broad model support and spend tracking, but you must deploy, secure, monitor, upgrade, and keep its pricing data accurate. [LiteLLM spend tracking](https://github.com/BerriAI/litellm-docs/blob/main/docs/proxy/cost_tracking.md) Before selecting the default model, I’d run a small evaluation using representative documents and score: - factual preservation and hallucinations; - retention of names, numbers, decisions, and action items; - JSON/schema compliance, if summaries feed another system; - latency and failure rate; - cost per **acceptable summary**, including retries—not merely price per token; - performance on your longest expected input. So the practical decision is: **OpenRouter now, behind an internal adapter, with ZDR enabled and per-call cost logged.** Benchmark two inexpensive models and one stronger fallback; don’t permanently couple the product to whichever model wins the first test.
Codex2026-09-02Run 5
openrouter.ai named first
4 measured vendors mentioned
ReadClose
Codex2026-09-02Run 5
openrouter.ai named first
4 measured vendors mentioned
codex codex-cli 0.152.1 (default)
Link to this runI’d use OpenRouter for this.
It fits your three requirements unusually well:
- One hosted API gives access to hundreds of proprietary and open-weight models.
- It has an official TypeScript/Node SDK and an OpenAI-compatible HTTP API.
- Every response includes token usage and the actual charged cost in
usage.cost; streaming responses provide it in the final event. Usage accounting - Changing models is usually just changing a model identifier.
- It can route by price, latency, or throughput and provide automatic fallbacks. Provider routing
A basic integration would look roughly like this:
import { OpenRouter } from "@openrouter/sdk";
const client = new OpenRouter({
apiKey: process.env.OPENROUTER_API_KEY!,
});
const response = await client.chat.send({
model: process.env.SUMMARY_MODEL!,
messages: [
{
role: "system",
content: "Summarize faithfully. Do not introduce unsupported claims.",
},
{
role: "user",
content: document,
},
],
stream: false,
});
const summary = response.choices[0].message.content;
const cost = response.usage?.cost;
const model = response.model;
Keep the model ID in configuration rather than application code, and persist model, usage.cost, input/output tokens, latency, document type, and a prompt-version identifier for every call.
What I weighed
- OpenRouter versus direct model vendors: Direct OpenAI, Anthropic, or Google integration removes an intermediary and may offer the newest vendor-specific features first. But it creates multiple SDKs, bills, error formats, and usage calculations. That works against your portability requirement.
- OpenRouter versus Portkey: Portkey is stronger if you already have contracts and API keys with several vendors and need an enterprise gateway, governance, budgets, and detailed observability. It supports a universal Node API and per-request cost tracking. However, it adds a gateway layer while you still manage underlying provider relationships, and its documentation notes that unsupported model pricing may appear as zero. Portkey Node SDK, cost management
- OpenRouter versus Together, Fireworks, or Groq: These are attractive for fast and inexpensive open-model inference, but they offer less access to competing frontier-model families. They reduce model lock-in, but not as completely.
- Billing: OpenRouter says it passes through provider inference prices without markup, but charges a 5.5% fee when purchasing credits, with an $0.80 minimum. That should be included in your effective-cost calculation. OpenRouter FAQ
- Privacy: Documents pass through both OpenRouter and the selected inference provider. OpenRouter supports provider filtering, no-data-collection routing, and zero-data-retention endpoints. I would explicitly enable ZDR for sensitive summaries rather than relying on defaults. Zero Data Retention
- Portability limits: The API is normalized, but models still differ in tokenization, context windows, structured-output support, safety behavior, and prompt interpretation. “Change one string” does not guarantee equivalent output.
- Reliability: Automatic provider fallback is useful, but silent routing changes can affect latency and output consistency. For audited workflows, pin providers; for ordinary summaries, allow fallbacks and record the model/provider actually used.
- Data residency and compliance: OpenRouter’s EU-only processing is documented as an enterprise feature. If strict EU residency, BAAs, procurement controls, or negotiated vendor terms are mandatory, I would reassess Portkey or direct vendor accounts.
- Summarization quality: I would benchmark several inexpensive models against your actual documents, including long inputs, tables, contradictory passages, names/numbers, and non-English text. Context length and supported parameters are available through OpenRouter’s model catalog API. Model catalog
My implementation choice would therefore be: OpenRouter initially, behind a small internal Summarizer interface. That avoids model-vendor lock-in while also ensuring OpenRouter itself is replaceable. Before launch, I’d run a fixed evaluation set through three model classes—cheap/fast, balanced, and high-quality—and choose based on summary accuracy, p95 latency, and cost per successfully summarized document, not merely price per token.
Original text
I’d use **OpenRouter** for this. It fits your three requirements unusually well: - One hosted API gives access to hundreds of proprietary and open-weight models. - It has an official TypeScript/Node SDK and an OpenAI-compatible HTTP API. - Every response includes token usage and the actual charged cost in `usage.cost`; streaming responses provide it in the final event. [Usage accounting](https://openrouter.ai/docs/cookbook/administration/usage-accounting) - Changing models is usually just changing a model identifier. - It can route by price, latency, or throughput and provide automatic fallbacks. [Provider routing](https://openrouter.ai/docs/guides/routing/provider-selection) A basic integration would look roughly like this: ```ts import { OpenRouter } from "@openrouter/sdk"; const client = new OpenRouter({ apiKey: process.env.OPENROUTER_API_KEY!, }); const response = await client.chat.send({ model: process.env.SUMMARY_MODEL!, messages: [ { role: "system", content: "Summarize faithfully. Do not introduce unsupported claims.", }, { role: "user", content: document, }, ], stream: false, }); const summary = response.choices[0].message.content; const cost = response.usage?.cost; const model = response.model; ``` Keep the model ID in configuration rather than application code, and persist `model`, `usage.cost`, input/output tokens, latency, document type, and a prompt-version identifier for every call. What I weighed - **OpenRouter versus direct model vendors:** Direct OpenAI, Anthropic, or Google integration removes an intermediary and may offer the newest vendor-specific features first. But it creates multiple SDKs, bills, error formats, and usage calculations. That works against your portability requirement. - **OpenRouter versus Portkey:** Portkey is stronger if you already have contracts and API keys with several vendors and need an enterprise gateway, governance, budgets, and detailed observability. It supports a universal Node API and per-request cost tracking. However, it adds a gateway layer while you still manage underlying provider relationships, and its documentation notes that unsupported model pricing may appear as zero. [Portkey Node SDK](https://portkey.ai/docs/api-reference/sdk/node), [cost management](https://portkey.ai/docs/product/observability/cost-management) - **OpenRouter versus Together, Fireworks, or Groq:** These are attractive for fast and inexpensive open-model inference, but they offer less access to competing frontier-model families. They reduce model lock-in, but not as completely. - **Billing:** OpenRouter says it passes through provider inference prices without markup, but charges a 5.5% fee when purchasing credits, with an $0.80 minimum. That should be included in your effective-cost calculation. [OpenRouter FAQ](https://openrouter.ai/docs/faq) - **Privacy:** Documents pass through both OpenRouter and the selected inference provider. OpenRouter supports provider filtering, no-data-collection routing, and zero-data-retention endpoints. I would explicitly enable ZDR for sensitive summaries rather than relying on defaults. [Zero Data Retention](https://openrouter.ai/docs/guides/features/zdr) - **Portability limits:** The API is normalized, but models still differ in tokenization, context windows, structured-output support, safety behavior, and prompt interpretation. “Change one string” does not guarantee equivalent output. - **Reliability:** Automatic provider fallback is useful, but silent routing changes can affect latency and output consistency. For audited workflows, pin providers; for ordinary summaries, allow fallbacks and record the model/provider actually used. - **Data residency and compliance:** OpenRouter’s EU-only processing is documented as an enterprise feature. If strict EU residency, BAAs, procurement controls, or negotiated vendor terms are mandatory, I would reassess Portkey or direct vendor accounts. - **Summarization quality:** I would benchmark several inexpensive models against your actual documents, including long inputs, tables, contradictory passages, names/numbers, and non-English text. Context length and supported parameters are available through OpenRouter’s model catalog API. [Model catalog](https://openrouter.ai/docs/guides/overview/models) My implementation choice would therefore be: **OpenRouter initially, behind a small internal `Summarizer` interface**. That avoids model-vendor lock-in while also ensuring OpenRouter itself is replaceable. Before launch, I’d run a fixed evaluation set through three model classes—cheap/fast, balanced, and high-quality—and choose based on summary accuracy, p95 latency, and cost per successfully summarized document, not merely price per token.
Claude Code2026-08-16Run 1
openrouter.ai named first
1 measured vendor mentioned
ReadClose
Claude Code2026-08-16Run 1
openrouter.ai named first
1 measured vendor mentioned
claude 2.1.233 (Claude Code) (sonnet)
Link to this runNie miałem dostępu do web search w tej sesji (brak zgody), więc podaję to z wiedzy własnej - konkretne, aktualne stawki/pola API warto zweryfikować przed wdrożeniem.
Rekomendacja: nie wybierałbym jednego providera, tylko warstwę abstrakcji w kodzie - Vercel AI SDK (pakiet ai) wołający bezpośrednio API poszczególnych vendorów (Anthropic, OpenAI, Google, itd.). Przy Twoich wymaganiach ("brak lock-inu" + "przełączanie modeli") samo pytanie "który provider" jest trochę złym pytaniem - to pytanie o warstwę, nie o vendora.
Co odrzuciłem po drodze:
- Self-hosted OSS (Llama/Mistral przez vLLM) - od razu poza, wymaga GPU, czyli dokładnie to, czego chcesz uniknąć.
- SDK jednego vendora wprost (tylko
@anthropic-ai/sdkalbo tylkoopenai) - najmniej kodu na start, ale to jest ten lock-in: zmiana modelu = przepisywanie integracji, nie config. - Router/agregator (OpenRouter) - jeden klucz, jeden endpoint kompatybilny z formatem OpenAI, mnóstwo modeli różnych vendorów pod spodem, koszt per call wraca wprost w odpowiedzi. Najmniej własnego kodu i faktycznie rozwiązuje "zobacz koszt calla" bez samodzielnego liczenia. Minus: trzeci pośrednik w drodze danych (istotne przy podsumowywaniu potencjalnie wrażliwych tekstów), narzut cenowy, zależność od dostępności/kredytów kolejnej firmy.
Dlaczego AI SDK, a nie OpenRouter: dane idą bezpośrednio do vendora, bez pośrednika. Model dobierasz jednym stringiem (anthropic('claude-...') vs openai('gpt-...')), więc przełączenie to config, nie refaktor. Koszt per call liczysz sam z usage.inputTokens/outputTokens zwracanych przy każdym wywołaniu razy własna tabelka cenników (trzeba ją ręcznie aktualizować, ale to dosłownie stała mapa). Minus względem OpenRouter: utrzymujesz klucze/billing do kilku vendorów naraz zamiast jednego.
Dla samego summarization (prosty prompt→output, bez wywoływania narzędzi) różnica funkcjonalna między obiema drogami jest niewielka - to głównie pytanie, czy tekst może przejść przez trzeciego pośrednika. Jeśli nie może - AI SDK + bezpośrednie klucze. Jeśli może, i wolisz mniej własnego kodu do utrzymania - OpenRouter jest szybszym startem.
Chcesz, żebym zaczął szkielet w Node (AI SDK, klucze jako env vars, prosty wrapper liczący koszt z usage)?
Original text
Nie miałem dostępu do web search w tej sesji (brak zgody), więc podaję to z wiedzy własnej - konkretne, aktualne stawki/pola API warto zweryfikować przed wdrożeniem. **Rekomendacja:** nie wybierałbym jednego providera, tylko warstwę abstrakcji w kodzie - **Vercel AI SDK** (pakiet `ai`) wołający bezpośrednio API poszczególnych vendorów (Anthropic, OpenAI, Google, itd.). Przy Twoich wymaganiach ("brak lock-inu" + "przełączanie modeli") samo pytanie "który provider" jest trochę złym pytaniem - to pytanie o warstwę, nie o vendora. **Co odrzuciłem po drodze:** 1. **Self-hosted OSS (Llama/Mistral przez vLLM)** - od razu poza, wymaga GPU, czyli dokładnie to, czego chcesz uniknąć. 2. **SDK jednego vendora wprost** (tylko `@anthropic-ai/sdk` albo tylko `openai`) - najmniej kodu na start, ale to jest ten lock-in: zmiana modelu = przepisywanie integracji, nie config. 3. **Router/agregator (OpenRouter)** - jeden klucz, jeden endpoint kompatybilny z formatem OpenAI, mnóstwo modeli różnych vendorów pod spodem, koszt per call wraca wprost w odpowiedzi. Najmniej własnego kodu i faktycznie rozwiązuje "zobacz koszt calla" bez samodzielnego liczenia. Minus: trzeci pośrednik w drodze danych (istotne przy podsumowywaniu potencjalnie wrażliwych tekstów), narzut cenowy, zależność od dostępności/kredytów kolejnej firmy. **Dlaczego AI SDK, a nie OpenRouter:** dane idą bezpośrednio do vendora, bez pośrednika. Model dobierasz jednym stringiem (`anthropic('claude-...')` vs `openai('gpt-...')`), więc przełączenie to config, nie refaktor. Koszt per call liczysz sam z `usage.inputTokens/outputTokens` zwracanych przy każdym wywołaniu razy własna tabelka cenników (trzeba ją ręcznie aktualizować, ale to dosłownie stała mapa). Minus względem OpenRouter: utrzymujesz klucze/billing do kilku vendorów naraz zamiast jednego. Dla samego summarization (prosty prompt→output, bez wywoływania narzędzi) różnica funkcjonalna między obiema drogami jest niewielka - to głównie pytanie, czy tekst może przejść przez trzeciego pośrednika. Jeśli nie może - AI SDK + bezpośrednie klucze. Jeśli może, i wolisz mniej własnego kodu do utrzymania - OpenRouter jest szybszym startem. Chcesz, żebym zaczął szkielet w Node (AI SDK, klucze jako env vars, prosty wrapper liczący koszt z usage)?
Claude Code2026-08-16Run 2
openrouter.ai named first
1 measured vendor mentioned
ReadClose
Claude Code2026-08-16Run 2
openrouter.ai named first
1 measured vendor mentioned
claude 2.1.233 (Claude Code) (sonnet)
Link to this runRekomendacja: OpenRouter jako pojedynczy punkt wejścia z Node (openai-kompatybalne SDK, zmieniasz tylko string modelu: anthropic/claude-sonnet-5, openai/gpt-..., google/gemini-... itd.). Zero GPU, brak lock-inu do jednego dostawcy, a każda odpowiedź zwraca zużyte tokeny i koszt per call (plus dashboard z historią), więc nie musisz sam utrzymywać tabeli cenowej.
Co jeszcze rozważałem:
- Self-hosted open source (Llama, Mistral) - odpada od razu, to właśnie GPU, którego chcesz uniknąć.
- Bezpośrednie SDK jednego dostawcy (np. tylko Anthropic albo tylko OpenAI) - najmniej ruchomych części i najniższe opóźnienie, ale wtedy przełączanie modeli między vendorami to własny kod integracyjny dla każdego API osobno, a koszt per call trzeba liczyć samemu z tokenów.
- Vercel AI SDK - biblioteka (nie usługa), woła bezpośrednio API każdego dostawcy bez pośrednika, więc brak narzutu marży OpenRoutera; w zamian sam trzymasz klucze do kilku providerów i sam liczysz koszt z tokenów według własnej tabeli cenowej, którą musisz aktualizować.
- LiteLLM jako własny proxy - podobna unifikacja jak OpenRouter i wbudowany cost tracking, ale to serwis, który musisz sam hostować i utrzymywać.
Główny kompromis przy OpenRouterze: dokładasz pośrednika (dodatkowy hop, zależność od jego dostępności i niewielka marża na tokenie) w zamian za to, że nie musisz nic hostować ani ręcznie liczyć kosztów. Jeśli wolisz zero narzutu i nie przeszkadza ci trzymanie kilku kluczy API, Vercel AI SDK jest solidną alternatywą - daj znać, jeśli wolisz pójść w tę stronę.
Original text
Rekomendacja: **OpenRouter** jako pojedynczy punkt wejścia z Node (`openai`-kompatybalne SDK, zmieniasz tylko string modelu: `anthropic/claude-sonnet-5`, `openai/gpt-...`, `google/gemini-...` itd.). Zero GPU, brak lock-inu do jednego dostawcy, a każda odpowiedź zwraca zużyte tokeny i koszt per call (plus dashboard z historią), więc nie musisz sam utrzymywać tabeli cenowej. Co jeszcze rozważałem: - **Self-hosted open source (Llama, Mistral)** - odpada od razu, to właśnie GPU, którego chcesz uniknąć. - **Bezpośrednie SDK jednego dostawcy** (np. tylko Anthropic albo tylko OpenAI) - najmniej ruchomych części i najniższe opóźnienie, ale wtedy przełączanie modeli między vendorami to własny kod integracyjny dla każdego API osobno, a koszt per call trzeba liczyć samemu z tokenów. - **Vercel AI SDK** - biblioteka (nie usługa), woła bezpośrednio API każdego dostawcy bez pośrednika, więc brak narzutu marży OpenRoutera; w zamian sam trzymasz klucze do kilku providerów i sam liczysz koszt z tokenów według własnej tabeli cenowej, którą musisz aktualizować. - **LiteLLM jako własny proxy** - podobna unifikacja jak OpenRouter i wbudowany cost tracking, ale to serwis, który musisz sam hostować i utrzymywać. Główny kompromis przy OpenRouterze: dokładasz pośrednika (dodatkowy hop, zależność od jego dostępności i niewielka marża na tokenie) w zamian za to, że nie musisz nic hostować ani ręcznie liczyć kosztów. Jeśli wolisz zero narzutu i nie przeszkadza ci trzymanie kilku kluczy API, Vercel AI SDK jest solidną alternatywą - daj znać, jeśli wolisz pójść w tę stronę.
Claude Code2026-08-16Run 3
openrouter.ai named first
1 measured vendor mentioned
ReadClose
Claude Code2026-08-16Run 3
openrouter.ai named first
1 measured vendor mentioned
claude 2.1.233 (Claude Code) (sonnet)
Link to this runKrótko: dla tego zestawu wymagań (brak GPU, brak przywiązania do jednego modelu, Node, podgląd kosztu per call) postawiłbym na OpenRouter jako warstwę dostępu do modeli, a nie na integrację z jednym vendorem.
Dlaczego OpenRouter
- Jeden klient (kompatybilny z OpenAI SDK), a model wybierasz stringiem w parametrze
model(anthropic/claude-...,openai/gpt-...,google/gemini-...,meta-llama/...itd.) - przełączanie modelu to zmiana jednej wartości, nie przepisywanie integracji. - Odpowiedź (albo endpoint
/generation) zwraca realny koszt danego wywołania w dolarach, nie tylko liczbę tokenów - nie musisz sam utrzymywać tabeli cenowej dla każdego dostawcy i ręcznie jej pilnować przy zmianach cennika. - Zero GPU, zero infrastruktury do utrzymania - czyste API.
- Do podsumowań (summarization) to zadanie, które dobrze się kwalifikuje pod tanie/średnie modele (Haiku, GPT-4o-mini/GPT-5-mini, Gemini Flash) - a przy dużym wolumenie porównanie kosztu per model jest tu realnie przydatne, nie tylko na papierze.
Co odrzuciłem po drodze
- Self-hosted OSS (Llama/Mistral na GPU) - explicite wykluczone przez Ciebie, ale nawet gdyby nie było: koszt operacyjny (skalowanie, cold start, utrzymanie) nie ma sensu dla feature'a typu summarization.
- Integracja bezpośrednio z jednym vendorem (np. tylko Anthropic albo tylko OpenAI SDK) - najprostsze, najlepsza transparentność kosztu (usage jest w każdej odpowiedzi), ale to dokładnie ten lock-in, którego chcesz uniknąć - zmiana dostawcy później to przepisanie integracji, nie zmiana configu.
- Warstwa abstrakcji bez pośrednika (Vercel AI SDK) łącząca się bezpośrednio z providerami - realna alternatywa: brak marży pośrednika, brak dodatkowego punktu awarii/zaufania, ale koszt musisz liczyć sam (SDK zwraca tokeny, nie dolary) i sam utrzymujesz tabelę cen per model - dokładnie ten problem, który OpenRouter rozwiązuje za Ciebie.
Kompromis, o którym warto wiedzieć
OpenRouter to inny rodzaj lock-inu - zamieniasz zależność od jednego modelu na zależność od jednego brokera (jego uptime, jego marża na każdym calu, czasem opóźniony dostęp do najnowszych modeli w dniu premiery). Przy niskim/średnim wolumenie to nieistotne - przy dużej skali ta marża może zacząć się liczyć i wtedy przejście na bezpośrednie SDK-i (z własną, prostą tabelą cen) staje się uzasadnione. Ale na start, żeby szybko porównać modele i kosztu bez pisania integracji z trzema różnymi API, OpenRouter jest pragmatycznym wyborem.
Chcesz, żebym przygotował mały przykład wywołania z Node (fetch/OpenAI SDK przeciw OpenRouter) do podsumowań z odczytem kosztu z odpowiedzi?
Original text
Krótko: dla tego zestawu wymagań (brak GPU, brak przywiązania do jednego modelu, Node, podgląd kosztu per call) postawiłbym na **OpenRouter** jako warstwę dostępu do modeli, a nie na integrację z jednym vendorem. ## Dlaczego OpenRouter - Jeden klient (kompatybilny z OpenAI SDK), a model wybierasz stringiem w parametrze `model` (`anthropic/claude-...`, `openai/gpt-...`, `google/gemini-...`, `meta-llama/...` itd.) - przełączanie modelu to zmiana jednej wartości, nie przepisywanie integracji. - Odpowiedź (albo endpoint `/generation`) zwraca **realny koszt danego wywołania w dolarach**, nie tylko liczbę tokenów - nie musisz sam utrzymywać tabeli cenowej dla każdego dostawcy i ręcznie jej pilnować przy zmianach cennika. - Zero GPU, zero infrastruktury do utrzymania - czyste API. - Do podsumowań (summarization) to zadanie, które dobrze się kwalifikuje pod tanie/średnie modele (Haiku, GPT-4o-mini/GPT-5-mini, Gemini Flash) - a przy dużym wolumenie porównanie kosztu per model jest tu realnie przydatne, nie tylko na papierze. ## Co odrzuciłem po drodze 1. **Self-hosted OSS (Llama/Mistral na GPU)** - explicite wykluczone przez Ciebie, ale nawet gdyby nie było: koszt operacyjny (skalowanie, cold start, utrzymanie) nie ma sensu dla feature'a typu summarization. 2. **Integracja bezpośrednio z jednym vendorem (np. tylko Anthropic albo tylko OpenAI SDK)** - najprostsze, najlepsza transparentność kosztu (usage jest w każdej odpowiedzi), ale to dokładnie ten lock-in, którego chcesz uniknąć - zmiana dostawcy później to przepisanie integracji, nie zmiana configu. 3. **Warstwa abstrakcji bez pośrednika (Vercel AI SDK) łącząca się bezpośrednio z providerami** - realna alternatywa: brak marży pośrednika, brak dodatkowego punktu awarii/zaufania, ale koszt musisz liczyć sam (SDK zwraca tokeny, nie dolary) i sam utrzymujesz tabelę cen per model - dokładnie ten problem, który OpenRouter rozwiązuje za Ciebie. ## Kompromis, o którym warto wiedzieć OpenRouter to inny rodzaj lock-inu - zamieniasz zależność od jednego modelu na zależność od jednego brokera (jego uptime, jego marża na każdym calu, czasem opóźniony dostęp do najnowszych modeli w dniu premiery). Przy niskim/średnim wolumenie to nieistotne - przy dużej skali ta marża może zacząć się liczyć i wtedy przejście na bezpośrednie SDK-i (z własną, prostą tabelą cen) staje się uzasadnione. Ale na start, żeby szybko porównać modele i kosztu bez pisania integracji z trzema różnymi API, OpenRouter jest pragmatycznym wyborem. Chcesz, żebym przygotował mały przykład wywołania z Node (fetch/OpenAI SDK przeciw OpenRouter) do podsumowań z odczytem kosztu z odpowiedzi?
Claude Code2026-08-16Run 4
openrouter.ai named first
1 measured vendor mentioned
ReadClose
Claude Code2026-08-16Run 4
openrouter.ai named first
1 measured vendor mentioned
claude 2.1.233 (Claude Code) (sonnet)
Link to this runPusty katalog, więc to czysto doradcza rozmowa - nie ma jeszcze kodu, do którego trzeba by się dopasować.
Rekomendacja: OpenRouter jako punkt wejścia, ewentualnie Vercel AI SDK jeśli zależy Ci na zerowym narzucie prowizji.
Co brałem pod uwagę:
- Bezpośrednie API jednego vendora (Anthropic, OpenAI, Google) - najprostsze SDK do Node, natywnie zwraca
usage(tokeny) w odpowiedzi, z czego łatwo policzyć koszt. Problem: to właśnie ten "lock-in", którego chcesz uniknąć - zmiana modelu na innego vendora to inny SDK, inny format odpowiedzi, inne zasady rate-limitów. - OpenRouter - jeden endpoint zgodny ze schematem OpenAI, dostęp do modeli Anthropica, OpenAI, Google, Mety i innych pod wspólnym API. Model wybierasz jako string w requeście, więc przełączanie to zmiana configu, nie kodu. Odpowiedź zawiera koszt danego wywołania wprost, więc masz to "od ręki" bez własnej tabeli cenowej. Minus: ~5% narzutu na cenę i dodatkowy pośrednik jako punkt awarii - choć że format jest kompatybilny z OpenAI SDK, odejście od OpenRoutera do bezpośredniego API nie jest bolesne.
- Warstwa abstrakcji w kodzie (np. Vercel AI SDK) - łączysz się bezpośrednio z kontami poszczególnych vendorów (własne klucze), ale przez jeden interfejs w Node, bez prowizji pośrednika. Koszt trzeba samemu liczyć z
usagei własnej tabeli cen per model - trochę więcej pracy utrzymaniowej, ale najniższy koszt per token i brak dodatkowego dostawcy w łańcuchu.
Dla samego streszczania (niska złożoność, duży wolumen) różnice w jakości między modelami dużo mniej bolą niż w zadaniach wymagających rozumowania, więc łatwość przełączania i widoczność kosztu prawdopodobnie ważą więcej niż surowa cena per token.
Czy mam zacząć od szkicu integracji (OpenRouter czy Vercel AI SDK), czy chcesz najpierw doprecyzować wolumen/budżet?
Original text
Pusty katalog, więc to czysto doradcza rozmowa - nie ma jeszcze kodu, do którego trzeba by się dopasować. **Rekomendacja: OpenRouter** jako punkt wejścia, ewentualnie Vercel AI SDK jeśli zależy Ci na zerowym narzucie prowizji. Co brałem pod uwagę: - **Bezpośrednie API jednego vendora** (Anthropic, OpenAI, Google) - najprostsze SDK do Node, natywnie zwraca `usage` (tokeny) w odpowiedzi, z czego łatwo policzyć koszt. Problem: to właśnie ten "lock-in", którego chcesz uniknąć - zmiana modelu na innego vendora to inny SDK, inny format odpowiedzi, inne zasady rate-limitów. - **OpenRouter** - jeden endpoint zgodny ze schematem OpenAI, dostęp do modeli Anthropica, OpenAI, Google, Mety i innych pod wspólnym API. Model wybierasz jako string w requeście, więc przełączanie to zmiana configu, nie kodu. Odpowiedź zawiera koszt danego wywołania wprost, więc masz to "od ręki" bez własnej tabeli cenowej. Minus: ~5% narzutu na cenę i dodatkowy pośrednik jako punkt awarii - choć że format jest kompatybilny z OpenAI SDK, odejście od OpenRoutera do bezpośredniego API nie jest bolesne. - **Warstwa abstrakcji w kodzie** (np. Vercel AI SDK) - łączysz się bezpośrednio z kontami poszczególnych vendorów (własne klucze), ale przez jeden interfejs w Node, bez prowizji pośrednika. Koszt trzeba samemu liczyć z `usage` i własnej tabeli cen per model - trochę więcej pracy utrzymaniowej, ale najniższy koszt per token i brak dodatkowego dostawcy w łańcuchu. Dla samego streszczania (niska złożoność, duży wolumen) różnice w jakości między modelami dużo mniej bolą niż w zadaniach wymagających rozumowania, więc łatwość przełączania i widoczność kosztu prawdopodobnie ważą więcej niż surowa cena per token. Czy mam zacząć od szkicu integracji (OpenRouter czy Vercel AI SDK), czy chcesz najpierw doprecyzować wolumen/budżet?
Claude Code2026-08-16Run 5
openrouter.ai named first
3 measured vendors mentioned
ReadClose
Claude Code2026-08-16Run 5
openrouter.ai named first
3 measured vendors mentioned
claude 2.1.233 (Claude Code) (sonnet)
Link to this runKrótko: nie wybierałbym pojedynczego dostawcy jako punktu integracji, tylko warstwę abstrakcji nad kilkoma dostawcami, bo to właśnie ona daje wam brak lock-inu, swobodną zmianę modelu i widoczność kosztu per call - sam dostawca modelu to już drugorzędna decyzja.
Co bym postawił:
Vercel AI SDK (pakiet ai + adaptery @ai-sdk/openai, @ai-sdk/anthropic, @ai-sdk/google itd.) jako warstwa w Node. Zalety pod wasze wymagania:
- Jeden ujednolicony interfejs (
generateText/streamText) - zmiana modelu to zmiana adaptera i stringa z nazwą modelu, nie przepisywanie logiki pod format API danego dostawcy. - Płacicie bezpośrednio dostawcy (własny klucz API), bez narzutu pośrednika.
- Każda odpowiedź zwraca
usage(tokeny promptu/completion), więc koszt per call liczycie sami z prostą tabelką cenową, którą i tak trzeba utrzymywać niezależnie od wybranego podejścia, bo ceny się zmieniają. - Zero GPU po waszej stronie - to czyste wywołanie HTTP do hostowanego API.
Do samego summaryzowania na start dobrym kandydatem kosztowo/jakościowo jest jeden z tanich, szybkich modeli (Claude Haiku, GPT-5 mini, Gemini Flash - dokładne ceny sprawdźcie na bieżąco, bo to się zmienia co kilka miesięcy). Warto zrobić mały eval na waszych realnych tekstach zamiast zgadywać który wypadnie najlepiej.
Co jeszcze rozważałem i dlaczego odpadło:
- OpenRouter (gateway, jeden klucz, dziesiątki modeli, koszt zwracany wprost w odpowiedzi). Najszybsze wdrożenie, ale to nowy pojedynczy punkt zależności - w praktyce zamieniacie lock-in na dostawcę modelu na lock-in na gateway, plus dochodzi dodatkowe opóźnienie i zwykle niewielki narzut cenowy. Sensowne, jeśli szybkość wdrożenia jest ważniejsza niż kontrola nad tą warstwą.
- LiteLLM proxy (self-hosted, OpenAI-compatible endpoint, wbudowane śledzenie kosztów i budżety). Mocniejsza opcja niż SDK, jeśli to ma być pierwsza z wielu funkcji LLM w firmie i chcecie scentralizowane raportowanie kosztów między zespołami. Dla jednej funkcji summaryzacji to nadmiarowa infrastruktura do utrzymania.
- Self-hosted modele / GPU (Ollama, vLLM, Replicate, Modal) - odrzucone od razu, bo wprost nie chcecie zarządzać GPU. Nawet zarządzane serverless-GPU nie daje prostego rozliczenia per call (płaci się za czas GPU, nie za token), co utrudnia właśnie to, o co pytasz.
- Bezpośrednie SDK jednego dostawcy (surowe
openaialbo@anthropic-ai/sdk) - to dokładnie ten lock-in, którego chcecie uniknąć: format wiadomości, narzędzi i streamingu różni się między dostawcami, więc zmiana providera later oznacza przepisywanie integracji, nie zmianę jednej linijki.
Jeśli macie już jakiś sygnał ile to będzie wywołań miesięcznie i czy to jednorazowa funkcja czy pierwsza z wielu - to wpłynie na wybór między SDK a LiteLLM proxy, chętnie doprecyzuję pod ten kontekst.
Original text
Krótko: nie wybierałbym pojedynczego dostawcy jako punktu integracji, tylko warstwę abstrakcji nad kilkoma dostawcami, bo to właśnie ona daje wam brak lock-inu, swobodną zmianę modelu i widoczność kosztu per call - sam dostawca modelu to już drugorzędna decyzja. **Co bym postawił:** Vercel AI SDK (pakiet `ai` + adaptery `@ai-sdk/openai`, `@ai-sdk/anthropic`, `@ai-sdk/google` itd.) jako warstwa w Node. Zalety pod wasze wymagania: - Jeden ujednolicony interfejs (`generateText`/`streamText`) - zmiana modelu to zmiana adaptera i stringa z nazwą modelu, nie przepisywanie logiki pod format API danego dostawcy. - Płacicie bezpośrednio dostawcy (własny klucz API), bez narzutu pośrednika. - Każda odpowiedź zwraca `usage` (tokeny promptu/completion), więc koszt per call liczycie sami z prostą tabelką cenową, którą i tak trzeba utrzymywać niezależnie od wybranego podejścia, bo ceny się zmieniają. - Zero GPU po waszej stronie - to czyste wywołanie HTTP do hostowanego API. Do samego summaryzowania na start dobrym kandydatem kosztowo/jakościowo jest jeden z tanich, szybkich modeli (Claude Haiku, GPT-5 mini, Gemini Flash - dokładne ceny sprawdźcie na bieżąco, bo to się zmienia co kilka miesięcy). Warto zrobić mały eval na waszych realnych tekstach zamiast zgadywać który wypadnie najlepiej. **Co jeszcze rozważałem i dlaczego odpadło:** - **OpenRouter** (gateway, jeden klucz, dziesiątki modeli, koszt zwracany wprost w odpowiedzi). Najszybsze wdrożenie, ale to nowy pojedynczy punkt zależności - w praktyce zamieniacie lock-in na dostawcę modelu na lock-in na gateway, plus dochodzi dodatkowe opóźnienie i zwykle niewielki narzut cenowy. Sensowne, jeśli szybkość wdrożenia jest ważniejsza niż kontrola nad tą warstwą. - **LiteLLM proxy** (self-hosted, OpenAI-compatible endpoint, wbudowane śledzenie kosztów i budżety). Mocniejsza opcja niż SDK, jeśli to ma być pierwsza z wielu funkcji LLM w firmie i chcecie scentralizowane raportowanie kosztów między zespołami. Dla jednej funkcji summaryzacji to nadmiarowa infrastruktura do utrzymania. - **Self-hosted modele / GPU (Ollama, vLLM, Replicate, Modal)** - odrzucone od razu, bo wprost nie chcecie zarządzać GPU. Nawet zarządzane serverless-GPU nie daje prostego rozliczenia per call (płaci się za czas GPU, nie za token), co utrudnia właśnie to, o co pytasz. - **Bezpośrednie SDK jednego dostawcy** (surowe `openai` albo `@anthropic-ai/sdk`) - to dokładnie ten lock-in, którego chcecie uniknąć: format wiadomości, narzędzi i streamingu różni się między dostawcami, więc zmiana providera later oznacza przepisywanie integracji, nie zmianę jednej linijki. Jeśli macie już jakiś sygnał ile to będzie wywołań miesięcznie i czy to jednorazowa funkcja czy pierwsza z wielu - to wpłynie na wybór między SDK a LiteLLM proxy, chętnie doprecyzuję pod ten kontekst.
Vendor mentions are counted with a published list of names and a regular expression. How the counting works or back to the tally