Skip to content
Fantástico Mundo de Jon
RSS

Dissecting an Ollama Modelfile to Stop the Model from Improvizing in Code

Ollama patterns are for conversation. For code, the task is to reduce randomness in sampling, write a dry SYSTEM, and save it in a derived model. Step-by-step with Qwen2.5-Coder.

In a hurry? Ask Claude for the TL;DR — it reads the page and summarises it.

en Machine translation by mistral-small3.2:24b, reviewed by the author. Read the Portuguese original

I install Qwen2.5-Coder, open it in Ollama and ask for a function. The model responds well. I ask again for the same function and another signature comes up, another variable name, sometimes a motivational comment in the middle. It’s not a major bug. It’s sampling with a chat-like appearance.

Ollama’s patterns were designed for conversation. High enough temperature to seem loose, repetition penalty to avoid getting stuck, short context to fit on any machine. For code, almost everything that matters is the opposite: reduce randomness and stop letting the model “creative”.

In this post I pull the Modelfile of an already installed Qwen2.5-Coder, dissect parameter by parameter, build a SYSTEM focused on code, create the derived model and point Aider to it.

What comes with the installed model

I start with a Qwen2.5-Coder in GGUF via Ollama. Any tag works for this exercise; what matters is having the model pulled beforehand. The command that extracts the recipe:

ollama show qwen2.5-coder:14b --modelfile > Modelfile

(no output in terminal — the content went to the Modelfile file)

Extracts the model's recipe to the current directory

The resulting file looks like this (the FROM points to the local blob on your machine; the hash changes due to quantization and tag):

# Modelfile generated by "ollama show"
# To build a new Modelfile based on this, replace FROM with:
# FROM qwen2.5-coder:14b

FROM /usr/share/ollama/.ollama/models/blobs/sha256-...
TEMPLATE """{{- if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
"""
PARAMETER stop "<|im_start|>"
PARAMETER stop "<|im_end|>"
PARAMETER temperature 0.7
SYSTEM """You are Qwen, created by Alibaba Cloud. You are a helpful assistant."""

Three quick observations. The TEMPLATE is Qwen’s chat template (tokens im_start / im_end); messing with it unnecessarily is the shortest way to break the model. The stop tokens prevent generation from invading the next turn. And the factory SYSTEM is generic for an assistant, not for code.

In several tags, Ollama still fits other PARAMETERs on top of global defaults. When it doesn’t fit, runtime defaults apply. This is the layer I want to control intentionally, instead of inheriting what’s left from a conversation preset.

temperature: the button that affects code the most

temperature rescales logits before softmax. At zero, the model always picks the token with the highest probability. As it increases, the distribution flattens and less probable tokens enter the game.

Ollama’s default is 0.8. Notice that the Modelfile I extracted above already brings temperature 0.7: 0.8 is the global runtime default, and this specific tag overrides it with 0.7. Both layers exist, and the top one applies. In chat, this makes the prose less mechanical. For code, 0.8 is an invitation to vary variable names in the middle of a refactor, reverse import order, and invent branches that the problem statement didn’t ask for.

What changes in the generated code: with high temperature you see diversity between one attempt and another. With low temperature, the same prompt tends to converge to the same skeleton. To complete a function, generate a test or apply a small diff, convergence is a feature.

Value I use for code: 0.1 for daily use, 0.0 when I want stable diff (patch, rename, repeat the same refactor). Rarely do I go above 0.2. Above that, in my usage, “creativity” appears as inconsistency, not as a better solution.

top_p: the probability mass cutoff

top_p (nucleus sampling) restricts selection to the most probable tokens until their combined probability reaches p. Ollama’s default is 0.9: quite a long tail remains.

In code, a long tail is where almost correct identifiers, blog-like comment words, and occasionally a swapped operator live. Lowering top_p shortens this tail.

What changes in practice: top_p 0.9 with temperature 0.8 is the conversation combo. top_p 0.9 with temperature 0.1 already improves a lot because temperature dominates. That’s why I leave top_p at its default 0.9: with temperature at 0.1 it has little to cut, and stacking two filters only makes it harder to know which one caused what. If I raise the temperature to 0.2 for more open tasks (sketching a new API), then I close top_p to 0.8.

top_k: the token count ceiling

top_k limits sampling to the k most probable tokens. Ollama’s default: 40.

It’s a blunt cut. Unlike top_p, it doesn’t look at accumulated mass: it looks at quantity. In a well-concentrated distribution (idiomatic code after a clear prompt), the first few tokens already carry almost everything and top_k barely appears. In a flatter distribution, it prevents the 200th token in line from entering due to noise.

For code I leave 40 or lower it to 30. I don’t think it’s the main lever. If temperature is already at 0.1, tweaking top_k too much is micromanagement. I use top_k as a safety net, not as fine-tuning of style.

min_p: the relative floor many people ignore

min_p discards tokens whose probability is less than min_p times the probability of the most likely token. Ollama’s default: 0.0 (disabled).

The idea is elegant: if the best token has a prob of 0.6 and min_p is 0.05, all below 0.03 fall out. At low temperature this hardly changes anything. At medium temperature, it cleans up the tail without the rigidity of top_k.

For code with temperature 0.1, I leave 0.05 without expecting miracles — it’s a belt and suspenders. If someone insists on running code with chat-like temperature (0.6–0.8), min_p at 0.05 or 0.1 helps more than praying. I prefer not to be in that situation.

repeat_penalty: the double-edged sword in code

repeat_penalty > 1 penalizes tokens that have already appeared, pushing the model away from what it just said. Ollama’s default: 1.1.

In prose, 1.1 reduces annoying loops. In code, repeating is not a failure: i, err, ctx, self, return, the same function name three times in a file, the same type in each signature. Penalizing repetition pushes the model to creative synonyms in the worst possible place — identifier and keyword.

What I’ve seen in practice: with 1.1 or more, the chance of changing a stable name to a variant (userId becoming user_id in the middle of a patch) and escaping from a legitimate construction that only has one idiomatic form increases. For code I use 1.0 (neutral effect) or at most 1.05 if some model enters a rare token loop. I don’t use 1.1 conversation in a code model.

num_predict: when to stop generating

num_predict is the ceiling of tokens in the output. Ollama’s default: -1 (no artificial ceiling beyond context).

For infinite chat it makes sense. For code, output without a ceiling on a vague prompt is monologue: the model explains, re-explains, generates an example, generates a variant of the example. In tools like Aider this becomes noise and sometimes cuts in the middle of a long file by hitting another limit along the way.

I use 2048 or 4096 depending on the task. Completing a function: 2048 solves most cases. Generating a larger test file: 4096. I don’t set 128k of prediction just because the model accepts it — predictable and short errs less than long and rambling.

num_ctx: the window that the default hides

num_ctx is the size of context in tokens. Ollama’s global default is 2048, and this is the number that most surprises people. The model card mentions 32k or 128k; the server starts with 2k if you don’t ask for something else.

In code, 2048 disappears quickly: system prompt, current file, neighboring snippets, Aider’s message, diff. With 2k the model “forgets” the top import and reinvents it. Or worse: silently truncates and responds on incomplete context.

For code I start with 8192 in daily use and go up to 16384 or 32768 when Aider sends a real multi-file task. The ceiling depends on VRAM and KV cache, not model marketing. The memory calculation is for another post; here the warning is just: don’t trust the 2048 default to work in a repository.

The SYSTEM I record in the model

No parameter replaces clear instruction. Qwen2.5-Coder’s factory SYSTEM is for a helpful assistant. I replace it with a dry block, in English (models in this family still obey system instructions better in English than in long Portuguese), focused on what breaks code in practice:

SYSTEM """You are a senior software engineer working in a local coding agent.
Prefer correct, minimal changes over rewrites.
Match the style, naming, and abstractions already present in the repository.
When editing, output only what was asked: no tutorial, no preamble, no recap.
If the request is to write code, write code. Put brief comments only where the intent is non-obvious.
Do not invent files, APIs, or config keys that were not given.
If requirements are ambiguous, state the assumption in one line and proceed.
Never wrap the whole answer in a markdown fence unless the user asked for a document.
Refuse to guess secrets, credentials, or private keys; use placeholders instead.
"""

Why each line is there:

  • senior… local coding agent: anchors the tone and context of use (Aider, editor, not sofa chat).
  • minimal changes over rewrites: the number one vice is rewriting the entire file when they asked for a six-line patch.
  • Match the style…: holds back the urge to “improve” names and standards in the middle of a fix.
  • output only what was asked: cuts prefaces like “Of course! Here’s a robust implementation…”.
  • write code / comments only where…: useful comment yes; line-by-line narration no.
  • Do not invent files, APIs…: hallucination of neighboring modules is the most expensive failure mode in an agent with real repo.
  • ambiguous → assumption in one line: prefers explicit progress to eternal question, without hiding the guess.
  • Never wrap the whole answer…: avoids giant fence that interferes with diff application.
  • Refuse to guess secrets…: obvious safety belt, and the local model still tries to “complete” key if you let it.

I don’t put a list of languages. I don’t put “you are world-class”. I don’t put rules I don’t want real enforcement. A too-long SYSTEM becomes noise that the model starts ignoring in the middle of context.

The final Modelfile and create

Putting everything together in a file that starts from the official tag (more portable than the blob path):

FROM qwen2.5-coder:14b

PARAMETER temperature 0.1
PARAMETER top_p 0.9
PARAMETER top_k 40
PARAMETER min_p 0.05
PARAMETER repeat_penalty 1.0
PARAMETER num_predict 4096
PARAMETER num_ctx 16384

SYSTEM """You are a senior software engineer working in a local coding agent.
Prefer correct, minimal changes over rewrites.
Match the style, naming, and abstractions already present in the repository.
When editing, output only what was asked: no tutorial, no preamble, no recap.
If the request is to write code, write code. Put brief comments only where the intent is non-obvious.
Do not invent files, APIs, or config keys that were not given.
If requirements are ambiguous, state the assumption in one line and proceed.
Never wrap the whole answer in a markdown fence unless the user asked for a document.
Refuse to guess secrets, credentials, or private keys; use placeholders instead.
"""

I create the derived model:

ollama create qwen-coder-dev -f Modelfile

transferring model data using existing layer sha256-… creating new layer sha256-… writing layer sha256-… success

Create the adjusted model

I check if the parameters took effect:

ollama show qwen-coder-dev --modelfile

FROM qwen2.5-coder:14b TEMPLATE “”“…”“” PARAMETER temperature 0.1 PARAMETER top_p 0.9 PARAMETER top_k 40 PARAMETER min_p 0.05 PARAMETER repeat_penalty 1.0 PARAMETER num_predict 4096 PARAMETER num_ctx 16384 SYSTEM “”“You are a senior software engineer…”“”

Check what was recorded

The TEMPLATE and inherited stop tokens from the base model remain there. I didn’t redeclare them on purpose: what I don’t touch, I don’t break. When testing in chat:

ollama run qwen-coder-dev

write a pure function in Go that returns the unique ints from a slice, preserving order

response: direct function, without preamble and without recap.

Smoke test — the question goes to Ollama's interactive prompt

If too much prose still comes up, the SYSTEM didn’t take (wrong tag, wrong file) or the client is sending another system on top. Aider does this frequently — it’s the next step.

Pointing Aider to the created model

Aider talks to Ollama via OpenAI-compatible API. With the derived model running:

export OLLAMA_API_BASE=http://127.0.0.1:11434
aider --model ollama/qwen-coder-dev

In recent versions of Aider, the ollama_chat/ prefix also appears; use what your version documents. The point is not the flag, it’s not pointing to the raw qwen2.5-coder:14b if you just spent time fine-tuning the derived model.

Two details that matter in real usage:

  1. Aider sends its own system and history. The Modelfile’s SYSTEM continues to apply as the base layer of the model in Ollama, but Aider stacks editor instructions on top. If something seems “deaf” to your SYSTEM, check what Aider is injecting before raising temperature again.

  2. Context. The Modelfile’s num_ctx 16384 defines the window of the model in Ollama. If Aider sends more files than fit, the conversation truncates. It’s better to be selective about what you add to the session than inflate num_ctx until the GPU cries.

About size: 7B is usable for autocomplete and small patch; 14B is my standard working point on a machine with room; 32B improves reasoning of gross refactor and costs real VRAM. QwQ-32B and DeepSeek-R1, which arrived in the 2025 window, are another conversation — long reasoning, not dry code completion. For loop with Aider I stick to Qwen2.5-Coder.

Recipe recorded in qwen-coder-dev

Base
qwen2.5-coder:14b
temperature
0.1
top_p
0.9
top_k
40
min_p
0.05
repeat_penalty
1.0
num_predict
4096
num_ctx
16384

The set I leave recorded

Ollama’s conversation default is not neutral: it pushes the model to sound loose and fit on small GPU. Code asks for a different posture. Side by side:

Parameter Ollama Default What I Use for Code
temperature 0.8 0.1
top_p 0.9 0.9
top_k 40 40
min_p 0.0 0.05
repeat_penalty 1.1 1.0
num_predict -1 4096
num_ctx 2048 16384
Middle column: default values declared in the official Ollama Modelfile documentation, verified in the repository as of March 2025. Right column: what I leave recorded — usage judgment, not measurement. The two columns don't have the same weight of evidence.

Outside the table is the SYSTEM: minimal, in engineer mode, prohibiting preface and invented API.

I don’t need a new preset every morning. I need a derived model with boring sampling, honest window, and system that doesn’t ask for creativity. The rest is Qwen2.5-Coder’s weight doing the quiet work it already knows how to do.

Record the Modelfile, create the tag, point the tool to it. If code still comes with plot, resist the urge to raise temperature: most of the time the right thing is to lower it.