Back to the blog
Open SourceAIOllama

Ollaya: an Ollama for open source decision models

October 02, 2026·6 min read·Diego Horvatti

How much have you paid in tokens for a giant model to answer "yes" or "no"? That question sums up why Ollaya caught my attention. The project describes itself as "Ollama for open-source, Jev-style decision models". The idea is a simple way to download and run open source decision models locally. Just like Ollama did for chat LLMs. The source is the official site: ollaya.dev.

I'll break down what it promises, where it fits in a real backend and what I'd test before trusting it.

What Ollaya is, in one sentence

Think about what Ollama did for you. Before it, running a local model meant downloading weights by hand, fighting Python dependencies and hoping your GPU would cooperate. After it, things looked like this:

ollama pull qwen2.5:3b
ollama run qwen2.5:3b

Two commands and a local HTTP API. That was Ollama's big win: it turned a research project into a dev tool.

Ollaya wants to repeat that experience for a different kind of model. The focus isn't chatting or writing text. It's deciding. A decision model takes an input and returns a choice: approve or block, which queue, which route, which action to take next.

About the term "Jev-style": it's the project's own label. Before you start repeating it, read the docs and figure out what exactly it means. New jargon is great for marketing and terrible for code review.

Why small decision models make sense

Most of the AI usage I see in production isn't chat. It's a decision dressed up as generation. Examples everyone has written at some point:

  • classify a support ticket as "billing", "bug" or "question";
  • decide whether a comment goes to moderation;
  • pick which tool an agent should call;
  • flag a transaction as suspicious.

In all of these cases, the useful output fits in a few bytes. Even so, a lot of people send the request to a frontier model, pay for thousands of reasoning tokens and wait a second or two for the answer. It's like hiring a surgeon to pull out a splinter.

A small, specialized model running on your machine or your server wins on four points:

  • Latency: no round trip to an external API.
  • Cost: the marginal cost becomes CPU or GPU you already pay for.
  • Privacy: customer data never leaves your infrastructure.
  • Predictability: output restricted to a closed set of options is easier to test.

When the answer fits in an enum, the model doesn't need to fit in a datacenter.

What this looks like in code today

For a concrete reference, here's how I handle local classification today with Ollama. The trick is using the format field with a JSON Schema, which forces the output to follow the structure:

type Queue = 'billing' | 'bug' | 'question'

async function classifyTicket(text: string): Promise<Queue> {
  const res = await fetch('http://localhost:11434/api/generate', {
    method: 'POST',
    body: JSON.stringify({
      model: 'qwen2.5:3b',
      prompt: `Classify the support ticket:\n\n${text}`,
      stream: false,
      format: {
        type: 'object',
        properties: { queue: { enum: ['billing', 'bug', 'question'] } },
        required: ['queue'],
      },
    }),
  })

  const { response } = await res.json()
  return JSON.parse(response).queue
}

It works. But look at what's going on: I take a general purpose language model and tie it down with a schema and a prompt until it behaves like a classifier. It's an LLM pretending to be a decision model.

That's exactly the gap Ollaya is trying to fill. Instead of taming a chat model, you run a model that was built to choose from day one. If the tool delivers an experience similar to Ollama's (download, run, call over HTTP), the change in your code is small. The endpoint changes, the model changes, and classifyTicket keeps the same signature.

That's the part that interests me most. If your decision logic is already isolated behind a typed function, swapping the engine underneath becomes an infrastructure detail.

What I'd test before going to production

A new open source AI project deserves curiosity and suspicion in equal measure. Before changing anything in my backend, I'd run through this checklist:

  1. Build your own evaluation set. Pick 200 or 300 real, already labeled cases from your domain. A generic benchmark tells you nothing about your tickets.
  2. Compare against what you already have. Run the same set on your current model (external API or local LLM) and on the model served by Ollaya. Measure accuracy, p95 latency and memory usage.
  3. Check the license of each model. The tool being open source doesn't guarantee that every set of weights it distributes allows commercial use. This catches a lot of people.
  4. Test how it behaves under doubt. What happens when the input is ambiguous? Does the model return some kind of confidence score? Without that, you can't build a decent fallback.
  5. See who maintains the project. Recent commits, answered issues, releases with a changelog. Abandoned infrastructure tooling turns into debt fast.

Item 4 matters most to me. Automated decisions without a confidence level are dangerous. The pattern I use is simple:

const { queue, confidence } = await decide(text)

if (confidence < 0.8) {
  return sendToHuman(text)
}

return route(queue)

If the model doesn't give you something like confidence, you're automating in the dark.

Where Ollaya won't solve your problem

It's worth being honest about the limits. A local runtime for decision models won't help if:

  • You have no data to evaluate with. Without a labeled set, you can't tell whether the small model is getting it right or just looks like it is.
  • The decision needs open-ended context. If the model has to read a 40-page contract and reason about it, a compact classifier isn't the right tool.
  • Your volume is low. At 50 calls a day, the external API bill is negligible. Running one more service can cost more in attention than it saves in money.
  • You run on pure serverless. A local model needs a long-lived process with reserved memory. In a function that spins up and dies on every request, that means a heavy cold start.

None of these are flaws in the project. Just remember that "running locally" is an architecture decision, not a free upgrade.

Is Ollaya worth keeping an eye on?

Yes. Not because it will change everything tomorrow. It's because it bets on a direction I think is right: use a model sized to the problem. The industry spent two years trying to solve everything with the biggest model available. Now the bill has arrived, in latency, in cost and in sensitive data leaving the company.

My honest take: most backends using AI today would get the same or better results with a small, specialized, well-evaluated model. What was missing was a way to run this kind of model as easily as ollama run runs LLMs. If Ollaya delivers that with good docs and clear licenses, it goes into my toolkit.

In the meantime, the practical tip works with or without Ollaya: isolate each AI decision behind a typed function, with a closed output and a fallback to a human. That way, swapping the engine becomes a detail. That's how I've been structuring the projects you'll find in my projects.

LinkedIn summary

How much have you paid in tokens for a giant model to answer "yes" or "no"?

Most of the AI I see in production isn't chat. It's a decision dressed up as generation: classify a ticket, moderate a comment, pick the next action.

Ollaya wants to be the Ollama of open source decision models. You download it, run it locally and call it over HTTP.

When the answer fits in an enum, the model doesn't need to fit in a datacenter.

Before trusting it, I'd test it with real data from my domain, check the license of each model and require a confidence score so doubtful cases go to a human.

I wrote the full checklist on the blog, plus where it falls short. The link is in the comments.

#ArtificialIntelligence #OpenSource #Backend #LLM #SoftwareDevelopment