Skip to content
Adam Cagle
All field notes

AI workflows · The complete walkthrough

WHICH AIFOR WHICH JOB?LET’S FIND OUT.

How to use OpenRouter to compare models, test real tasks, and build a workflow around what your business actually needs.

By Adam Cagle14 min read

Ask five people which AI your business should use and you may get five very confident answers.

Ask what you need it to do, and now we can have a useful conversation.

Writing a customer email, finding a fact in a document, generating a video, and deciding where a support ticket belongs are different jobs. Giving all of them to the same model is convenient. It may also mean paying too much for the easy work and getting disappointing results on the hard work.

This is why I use OpenRouter. It gives me a place to explore models, compare their outputs, and build a working shortlist before connecting them into a business workflow.

In this episode of Super Intelligence, I walk through the catalog, test media models, look at benchmarks, create a limited API key, and explore Jev. This written guide adds a repeatable first experiment and explains how those experiments become something a team can use.

You do not need to test every model. You have a business to run. Possibly even a life.

What OpenRouter actually does

OpenRouter provides access to models from multiple AI companies through a shared platform and API. You can experiment in its browser interface, then connect compatible software using an API key. Its quickstart explains the integration options.

An API is a way for software to request work from another service. Your application sends a request; a model produces a response; your application decides what happens next.

OpenRouter makes that connection easier to reuse across models. It does not automatically know your business rules, connect your CRM, or decide which answers your customers should receive. Those are parts of the workflow we design around it.

The model and the app are different things, too. Accessing a model through an API does not reproduce every feature of the consumer chat product built around that model. Memory, document search, permissions, and tool access need to be provided by the application where appropriate.

Start with one job, not a favorite model

Before opening the catalog, finish this sentence:

“When we receive [input], we need [output], and a good result must [requirements].”

For example: “When we receive a customer inquiry, we need a draft reply, and a good result must use our current service information without inventing a price or promising availability.”

That is a testable job. “Make our company more AI” is a meeting that should have been an email.

Write down the input, the required output, and what would make an answer unacceptable. Include who approves the work and how quickly it needs to be ready. Now the model comparison has a purpose.

If your real process crosses several systems or nobody agrees where it starts and ends, bring that process to a conversation with me. Mapping the work is often the most useful first step.

Step 1: Set up an account and a small test budget

  1. Visit openrouter.ai and create an account, or sign in.
  2. Open the account’s Credits area. Review the purchase amount and checkout fees before adding funds.
  3. Start with a budget you are comfortable spending on an experiment. In the video, I describe starting with $20. That is my example, not a requirement or a promise about how much work it buys.
  4. Keep automatic replenishment off while you learn how your tests consume credits.
  5. Find the Activity area so you know where to check actual usage afterward.

Browser chat and API requests use the account’s credits. Model rates differ, and media generation can have different billing units from text. OpenRouter describes its billing and credit-purchase fees in the official FAQ.

Free models can be useful for exploration, but check their limits and data policies. Free is a price, not a quality certification.

Step 2: Read the model catalog like a shopping list

Open the model catalog and filter for the job you actually have. The catalog includes text, image, video, audio, speech, transcription, embeddings, rerankers, and decision models. Availability changes, so the counts you see in the recording are a snapshot. Model categories and capabilities

Here is what those categories mean in practical terms:

What you need What to investigate What to check
Draft or analyze text Text and reasoning models Accuracy, instructions, tone, and usable output
Create an image Image generation models Composition, consistency, text rendering, and editing control
Produce a video shot Video models Motion, continuity, duration, reference support, and cost per usable shot
Read a script aloud Speech models Pronunciation, delivery, voice fit, and generation time
Transcribe a recording Transcription models Names, terminology, speaker handling, and accuracy
Find relevant material Embeddings and rerankers Whether the right source is retrieved for real questions
Make a bounded judgment Decision models Whether the selected label or action meets your criteria

Open an individual model page and look beyond the description. Check supported inputs and outputs, pricing, context length, available providers, and the controls the model supports. A video model that cannot use your reference image is not a bargain for a job that depends on that image.

Context is working space, not automatic memory

A context window limits how much material the model can work with in a request. Input, conversation history, tool results, and generated output can consume that budget; output limits may also be smaller than the headline context number.

For a typical chat-completions integration, your application supplies the relevant conversation and source material. You do not have to resend your entire company history every time someone asks about opening hours. You do need a deliberate way to select and supply the right information.

Price is more than the number beside the model

For text, compare both input and output rates. A cheap input rate can still produce an expensive run if the model generates long answers, makes repeated attempts, or requires lots of human correction.

My useful number is cost per accepted result: what did it cost to get something we could actually use? Include retries and review effort in that discussion. A bargain that needs to be rewritten three times is developing an expensive personality.

Step 3: Compare models on the same task

Use OpenRouter Chat for a first text comparison. Select a few candidate models in the model picker and use its comparison controls where available. You can also run separate conversations with the same instructions and input. For image, video, or speech, use the playground exposed on the relevant model page and check the cost before generating.

Start with three candidates: a lower-cost option, a more capable reference option, and another credible model suited to the task. This is a shortlist, not a commitment.

For the first pass, keep the source material, instructions, requested format, and output length comparable. Start fresh conversations so one model does not benefit from context the others never received. Later, you can tune prompts for each model and compare the finished approaches.

Here is a fictional example you can copy without exposing client information:

You are drafting a customer reply for Sample Studio.

Approved facts: We offer headshot sessions and product photography. A team member must confirm dates and pricing. We have no approved price list or live calendar in this task.

Customer message: “Can you photograph eight products next Thursday? What would it cost, and can I get the images the same day?”

Write a friendly reply of no more than 100 words. Acknowledge the request, ask for the missing details needed to quote, and identify what a team member must confirm. Do not invent prices, availability, turnaround times, or a booking confirmation. Return the draft reply followed by a short list titled “Needs confirmation.”

Read the responses. Did any model promise next Thursday? Invent a price? Quietly turn a request into a confirmed booking? Those failures matter more than which answer had the nicest adjectives.

Then try variations: a vague request, conflicting details, an unsupported service, and a customer insisting on an immediate guarantee. A small set can expose obvious problems. It is not enough to establish production reliability.

Step 4: Use benchmarks for clues, then keep score

In the video, I look at media benchmarks and rankings to find candidates. Side-by-side outputs are especially useful when the work is visual or audible. You can hear a flat delivery or see a broken detail without needing a leaderboard to explain it.

Still, inspect what a ranking measures. Popularity or token usage tells you what people are using. It does not establish which model is best for your work. A benchmark measures performance on its test, which may be very different from your customer’s request.

Keep a simple scorecard for your own evaluation:

Measure What to record
Correctness Did it use the supplied facts and avoid invented details?
Requirements Did it follow the format, limits, and business rules?
Usability Could a person use this result without substantial repair?
Speed How long until the complete, usable result arrived?
Cost What was charged, including failed attempts and retries?
Failure behavior Did it expose missing information or confidently fill the gap?

For speech, include your business name and difficult terms in the test. For video, evaluate the whole shot, not only its attractive first frame. For document work, include messy source material as well as the beautifully formatted example.

Set the quality threshold first. Then compare price and speed among the options that pass it. Record the model, provider, settings, and test date so you can repeat the evaluation when something changes.

Step 5: Create an API key without giving away the wallet

You can explore in the browser first. Create a key when you are ready to connect an application or a tool that supports OpenRouter.

  1. Open Dashboard → API Keys, or visit API Keys.
  2. Choose Create Key or New Key, depending on the current interface.
  3. Give it a descriptive name, such as customer-reply-test.
  4. Set a small credit limit and an expiration appropriate to the experiment. Check whether the limit is lifetime or resets on a schedule.
  5. Copy the secret when it is shown and save it in your approved secret store. If you lose it, replace the key rather than trying to recover it from a screenshot.
  6. Test one request, then inspect Activity before running a larger batch.

A key’s spending limit is a cap on access to account funds. It does not move money into a separate wallet. Naming a key also does not automatically create an environment variable with that name. OpenRouter authentication

Keep the secret out of chat messages, public repositories, browser code, screenshots, and logs. Store application keys on the server, using a secret manager or the hosting platform’s protected environment variables. If a key is exposed, revoke it and replace it.

An environment variable is a named value your program can read while it runs. OPENROUTER_API_KEY is a common name; the value is the secret. The program needs an explicit connection to that value. An assistant does not gain secure access just because you tell it where the key lives, and software with broad execution access may be able to read it. Give the tool only the access it needs.

Step 6: Connect the model to your actual workflow

For compatible text-chat integrations, the usual connection details are:

Setting Value or action
API base URL https://openrouter.ai/api/v1
API key The secret supplied through your server’s protected configuration
Model The exact model ID copied from its OpenRouter page
Output requirements The instructions, format, and limits for this particular task

OpenRouter’s integration quickstart includes examples. Image, video, audio, and decision APIs can require different request formats, so check the documentation for the selected capability. Changing a model name does not make incompatible features interchangeable.

If an AI coding assistant is helping, give it the task and configuration names, not the secret. Here is a useful starting brief:

Build a server-side OpenRouter integration for this task: [describe the task]. The runtime will receive a secret named OPENROUTER_API_KEY. Never display or log its value, return it to the browser, or include it in source control.

Use the current official documentation and the exact model ID I select. Validate the response against these requirements: [list requirements]. Put limits on requests, output length, retries, and spending. Keep external actions disabled until a person approves them. Show me the test results, errors, and estimated operating cost before release.

For a customer-reply workflow, the shape might be: receive inquiry → find approved information → draft reply → check facts and requirements → request approval → send through the existing system.

Some of those steps need a model. Others need ordinary code. Dates, totals, permissions, and required fields are often better checked directly than debated by an AI committee.

For a talking assistant, retrieval, answer drafting, and speech can each use different tools. Approved recordings can cover common questions without generating new speech every time, although storage and delivery still have operating costs.

This is the part I build for clients: the connections, routing, checks, and approvals around the model. You can explore Switchboard for a guided example of task-based routing, or talk with me about your own workflow.

Gates and fallbacks: useful answers need a stopping point

A gate is a check that determines whether work can continue. Required fields are present. A cited source exists. The output uses an allowed category. The person approving a change has the right permission.

OpenRouter supports structured outputs on compatible models. That helps software receive a predictable shape, but valid JSON is not proof that the answer is true.

Similarly, model fallbacks can try another model when a request encounters an error. An answer that arrives successfully but contains a wrong fact needs your own quality checks. Availability and correctness are different problems.

My design goal is to detect defined failures, limit repair attempts, and hand unresolved cases to a person. Tests help us understand what those checks catch and what they miss. No gate makes an AI infallible.

Batch and Jev: two features worth understanding

Batch is for work that can wait

Preparing tomorrow’s document summaries is different from answering a customer who is waiting now. Batch processing groups requests for asynchronous completion. OpenRouter’s Batch API announcement describes a 24-hour completion window and discounted rates on supported models. Check the specific model and workload restrictions before planning around it.

A backlog of approved documents or an evaluation set can be a good candidate. A live voice conversation generally is not. The practical question is whether your work needs immediate delivery or simply needs to be ready by a deadline.

Jev makes decisions; Jev Router chooses a model

These names are related, but the jobs differ. Jev is a structured decision model from TypeSafe. It returns typed judgments and probabilities that software can use for classification and routing. Your application decides how those judgments affect the next action.

Jev Router uses Jev to choose a model and reasoning effort for a request. Its pricing is tied to the model that serves the request. Do not assume the decision model’s displayed output price means the routed answer is free.

Both are worth testing against your own examples. A router’s recommendation still needs to satisfy your quality, privacy, speed, and cost requirements. It is another candidate in the evaluation, not a substitute for having one.

Should you stay on OpenRouter or go direct?

I like OpenRouter for exploration because it makes comparison convenient. I also evaluate direct provider connections when a workflow’s requirements justify them.

But direct is not automatically faster or cheaper, and OpenRouter can be used in production. Compare the actual route under your workload: total latency, provider features, privacy requirements, operating cost, and the work involved in maintaining multiple integrations. OpenRouter’s performance guide explains several factors that affect response time.

Caching is not exclusive to direct connections. OpenRouter supports prompt caching for supported models and providers, and it now offers batch processing too. Choose based on the features and economics you can actually measure.

Before sending client material, also review the chosen providers’ logging and retention policies. Start experiments with fictional or appropriately redacted inputs. Privacy requirements apply to fallback routes as well as the first model you select.

Whichever connection you choose, rerun the same evaluation after changing it. “We only changed the provider” still means something changed.

Your first useful experiment

Pick one recurring task. Define an acceptable result. Try a few models on the same inputs. Record the mistakes, time, and cost. Then decide whether the result is worth developing further.

That is a productive afternoon. Collecting twenty impressive screenshots without knowing what you would use them for is a different hobby.

If you want help turning the experiment into a system your team can depend on, let’s talk about what you need built. Tell me what arrives, what should come out, which tools you already use, and where people need to approve the work. We can start there.

I work with businesses on custom AI workflows, agents, and implementation, remotely or on site. See consulting options, or email adam@adamcagle.com.

Live long and prosper. And test the expensive model before you give it every job in the company.

Platform details checked October 7, 2026. Model availability, interfaces, limits, and pricing can change; the official links throughout this guide are the current reference points.