AI and Testing: LangChain Templates

In this post I’m going to follow the thread from the previous post and dig more into the LangChain ecosystem and start looking at the idea of templates for prompts.

Before we get going, there are two bits of housekeeping to attend to.

Continue with LangSmith?

In the previous post, I had you connect to LangSmith. Do you need to do that for all of these posts? No, you don’t. I used LangSmith to show you how to run your scripts against an observability platform. You won’t really need LangSmith for a lot of these initial posts and thus I’m not going to focus on that as much, although I will come back to it.

If you do want to use LangSmith, and assuming you followed the .env file setup from the previous post, just make sure you include these lines in your code:

Another way to do it is keep those lines in your code at all times but change the LANGSMITH_TRACING in your .env file to false. Only set it back to true when you want tracing.

Type Errors

A challenge you often find in this context is the tension between Python’s static type checking and the highly dynamic nature of certain frameworks, like LangChain. While certain code behaves entirely correctly at runtime, static analyzers can’t always determine the exact types of values flowing through the system or being returned from the API and thus flag the code as being problematic.

Frameworks that rely heavily on composition, late binding, and generic abstractions, such as LangChain, intentionally defer many type decisions until execution time, making precise static typing impractical.

In editors like Visual Studio Code, this shows up as frequent Pylance warnings (usually of the “partially unknown type” variety) that would otherwise require widespread use of # type: ignore comments. Instead of suppressing individual warnings inline like this, a better approach for my pedagogical posts here is to provide you with a configuration that relaxes specific type-checking rules globally.

The easiest way to do that is to create a pyrightconfig.json file in your project directory and put the following in it:

I realize that from a quality perspective, this can seem questionable. I would say that this configuration reflects a conscious tradeoff and that in pedagogical contexts like mine, this helps readers focus on behavior, design, and reasoning first rather than why their IDE keeps complaining about things that obviously work.

Hitting Models Directly

Let’s take a look at hitting the model directly without using LangChain at all.

This code sends a direct HTTP request to Ollama running on your machine. Think of it like filling out a form and mailing it to get a response: the payload is your form with the model name and your question. The stream: False tells Ollama to send the complete answer all at once rather than word-by-word. We then dig into the response JSON to extract just the text answer we care about.

Now, let’s take a look at the exact same logic using just what we learned in the first post:

Both code examples do the same thing: send a message to Ollama and get back the response. LangChain acts like a universal translator and toolkit: it wraps the messy details in a simple, consistent way. What it adds specifically is a cleaner API (no manual JSON payload construction), automatic streaming handling, a consistent interface across different LLM providers (OpenAI, Anthropic, etc.), and, crucially, integration with the LangChain ecosystem: messages, chains, agents, prompt templates, and so on.

It’s part of that ecosystem that I want to talk about here: prompt templates. That will lead us into the chains part. However, before I get into that, let’s do one more variation on this that uses LangChain messages.

What you can see here is that we shifted abstractions a bit

  • HTTP + JSON + roles + transport
  • ChatOllama.invoke(“string”)
  • ChatOllama.invoke([HumanMessage(…)])

Each step introduces exactly one new idea and you will find that’s a lot of what happens when learning this stuff. You will be shifting between abstractions. Granted, you could argue that’s the case with any code-based implementations at all. This is my way of saying just make sure you are paying attention to the sometimes subtle differences in code.

Now let’s jump into one of those LangChain ecosystem elements.

Templating Our Prompts

Prompt templates let you create reusable question patterns, sort of like a fill-in-the-blank form you can use repeatedly with different values. Let’s get some initial logic in place to see how this all works.

Here, we’re using a ChatPromptTemplate, which is essentially a structured script that helps you organize a conversation between a human and an AI. With this script we’re not actually running anything against a model yet. We’re just examining what LangChain created and how it organized this information. You see should be this:


input_variables=['planet'] input_types={} partial_variables={} messages=[HumanMessagePromptTemplate(prompt=PromptTemplate(input_variables=['planet'], input_types={}, partial_variables={}, template='How many of planet Earth could fit inside {planet}?'), additional_kwargs={})]

You can see that the logic has recognized that your {planet} is an input variable, meaning this is something that will be expected to be substituted into the prompt.

Considering a bit more of that output, the nested structure of HumanMessagePromptTemplate and PromptTemplate exists because chat models work with roles. They need to know whether a message is from a human user, the AI assistant, or a system instruction. ChatPromptTemplate automatically wraps your prompt with this role information, in this case, marking it as coming from a human.

What you really have here is a structure like this:


ChatPromptTemplate(
  input_variables=['planet'],
  messages=[
    HumanMessagePromptTemplate(
      prompt=PromptTemplate(
        template='How many of planet Earth could fit inside {planet}?'
      )
    )
  ]
)

Effectively, a PromptTemplate formats text; a ChatPromptTemplate assembles role-aware message templates into a structured conversation that can be rendered into chat messages.

Let’s add a bit more code:

If you run this, you’ll get the following for the print statement we added:


Human: How many of planet Earth could fit inside Jupiter?

Let’s break down what just happened with two key methods.

  • The .from_template() method creates the template. It takes your string with {planet} and turns it into a reusable PromptTemplate object that LangChain understands.
  • The .format() method uses that template. It fills in the blank: wherever we had {planet}, it now says “Jupiter”. This is how you turn your template into an actual prompt.

Together, these two steps let you: (1) define a pattern once, then (2) reuse it with different values.

Execute Templates Against the Model

Wonderful, so templating appears to be working. I trust you can see that the templating is fairly simple. With this idea in place, let’s put in some of our logic from the previous post (and similar to what we did earlier in this post) to make this example executable against the model.

The key line there is the call to model.invoke(). That’s what we used in the previous post as well. The invoke method is what executes your prompt against the model. What this script does is essentially replicate what we did in the previous post (a simple, static prompt) but with the addition of the templated (dynamic) part, and thus building on what we started with in this post.

Further Parameterizing a Prompt

What about some of those other values that we saw in the data structure? The input_types refer to optional type specifications for your input variables. Those were empty because I didn’t specify any types. The partial_variables refer to pre-filled variables that are already set in the template. Again, these were empty because I haven’t pre-filled anything.

Let’s consider a different prompt to show all this in action.

Here we’re specifying that particle should be a string (input_types), and we’re pre-filling units with “kilograms” (partial_variables). This means you can query different particles without re-specifying the units each time. The units are locked in, as it were, to the template.

Some alternative units for variation would be “electron volts (eV)”, which is much more appropriate for particle physics, or “atomic mass units (amu)”.

Being Runnable

One thing I’ve been doing here is using the format approach with ChatPromptTemplate. In other words, we’ve used logic like this:

Let’s change line 20 in our previous script so it looks like this:

The key difference is that .invoke() is part of LangChain’s Runnable interface, which means this template can now be chained with other components. This means you can connect it directly to models, parsers, or other prompts, creating a pipeline where the output of one step automatically feeds into the next. Think of it like connecting LEGO blocks instead of manually passing data between separate pieces.

Also, when you print the full prompt with this change, you’ll notice .invoke() doesn’t return a plain string like .format() did. You’ll see this:


messages=[HumanMessage(content='What is the mass of the electron? Answer in kilograms.', additional_kwargs={}, response_metadata={})]

What gets returned is a LangChain message object that other Runnable components can work with directly.

Looking at the Thinking

Something worth calling out here is that the LLM, at least if you’re using Qwen3, is doing a lot of “thinking” when it responds to this prompt. However, you’re not actually seeing that in the content that is returned. There’s a way to check the thinking: use the model REPL. Go into the Qwen3 model via Ollama:

  ollama run qwen3

Now enter in the prompt: How many of planet Earth could fit inside Jupiter?

You might get something like this (which you don’t have to sit and read; I’m just providing this to give you the flavor):


Thinking...
Okay, so I need to figure out how many Earths can fit inside Jupiter. Hmm, let's start by recalling some basic facts about the sizes of these planets. I remember that Jupiter is the largest planet in our solar system, so it's definitely bigger than Earth. But how much bigger exactly?

First, I should find the volumes of both planets because volume is a three-dimensional measure, which is more accurate for comparing how many Earths can fit inside Jupiter. The volume of a sphere is given by the formula (4/3)?r³, right? So if I can get the radii of both Jupiter and Earth, I can calculate their volumes and then divide Jupiter's volume by Earth's volume to find out how many Earths would fit inside.

Wait, do I remember the exact radii? Let me think. I think Earth's radius is about 6,371 kilometers. For Jupiter, I believe it's much larger. I recall that Jupiter's radius is approximately 69,911 kilometers. Let me check if that's correct. Yes, I think that's right. So Jupiter's radius is roughly 11 times Earth's radius. Wait, 69,911 divided by 6,371... Let me do that division. 6,371 times 11 is 70,081, which is close to 69,911. So maybe it's about 11 times? But actually, the exact value might be a bit less. Let me calculate it precisely.

So, 69,911 divided by 6,371. Let's see, 6,371 * 10 = 63,710. Subtract that from 69,911: 69,911 - 63,710 = 6,201. Then, 6,201 divided by 6,371 is approximately 0.97. So total is 10.97, which is roughly 11 times. So Jupiter's radius is about 11 times Earth's radius.

But wait, I think the exact value might be a bit different. Maybe I should look up the exact radii. Wait, but since I can't actually look things up, I need to rely on my memory. I think the exact radius of Jupiter is about 69,911 km, and Earth is 6,371 km. So the ratio is 69,911 / 6,371 ? 11. So that's the radius ratio.
....

Notice how the model is entering an extended reasoning/thinking mode, doing verbose calculations and self-corrections, but eventually arriving at an answer. It’s that answer that you get returned in the content key of our response variable in the script.

This goes back to the rethinking of models I mentioned in the previous post. Qwen3 does a lot of this chain-of-thought. Other models do not. It’s instructive to try different models and see what they do.

The reason you don’t get all that thinking in the response is because LangChain’s invoke functionality on the ChatOllama object waits for the complete response and only returns the final result, filtering out intermediate thinking tokens. To be sure, the model is still doing all that thinking internally; you just don’t see it.

This is why the model sometimes appears to “hang,” and you hear your computer fan spin up or your GPU card kick into high gear and notch your electric bill up a bit. The model is actually generating thousands of tokens of reasoning that you can’t see, sometimes sitting there second-guessing or correcting itself.

This is a key testing point! What you’re seeing here is intentional design: most API consumers want the final answer, not the reasoning process. Yet, it’s the very reasoning process that we often want to reason about in our testing!

This is a major blindspot not just for testing but also for debugging and optimization!

Models like Qwen3 that have built-in chain-of-thought reasoning are doing substantial hidden work. Without visibility into that, you can’t (as easily) tune prompts to reduce unnecessary reasoning, understand why certain phrasings trigger verbose thinking, or optimize for speed and quality trade-offs.

I should also note that LangSmith would not help you with this because LangSmith is only dealing with what is returned as the answer, not the reasoning behind the answer. Put another way, LangSmith can only deal with what LangChain provides to it.

I’m not going to belabor this point too much here but I did want to at least bring this up because it’s one of the areas you need to be aware of when you consider testing a model. When I get this series to dealing with tools like DeepEval or Ragas, we’ll deal with how to look at this hidden work.

Structured Conversations

Let’s try a variation on what we’ve been doing to make those messages we looked at earlier a little more front-and-center.

This code demonstrates how to create reusable, structured prompts using LangChain’s message template system. This time we’re building a conversation framework; a structure that includes both system instructions and human questions, reusable with different inputs.

This introduces two new components working together. The SystemMessagePromptTemplate sets the AI’s role and behavior. Think of it as “behind-the-scenes instructions” that shape responses. Here we’re telling the model to act as a relativity expert for non-specialists. The HumanMessagePromptTemplate, which we briefly saw earlier, creates the user’s question template with the {concept} placeholder, just like we did with {planet} and {particle} earlier.

We combine both templates into a ChatPromptTemplate (system + human), then invoke it with “time dilation” to produce the actual message sent to the model.

The Value of Templates

Why use templates instead of simple strings?

First, reusability and scalability, both internal qualities. You can change “time dilation” to, say, “length contraction” without rewriting the entire prompt. In real applications, you can swap variables programmatically across hundreds of queries.

Second, separation of concerns. System instructions stay separate from user questions. This matters because you can update one without touching the other. For example, you could modify the expertise level without changing the question structure, or vice versa.

Third, consistency. The AI maintains the same persona (“expert in relativity”) across multiple queries without you having to repeat the system prompt every single time.

This pattern becomes essential when building chatbots, RAG systems, or any application where you need consistent, structured interactions with an LLM. That pattern being in place is what enables various forms of testing.

There actually is a more compact way to do that above script that many people prefer and it highlights the message sequence structure a bit better.

Does that look familiar? It’s very much like the script we ended the previous post with, in terms of “system” and “human” messages. This latest iteration of the script does exactly the same thing as the prior one, but more compactly.

LangChain lets you specify messages as Python tuples, which are paired values in parentheses. In this case the tuple is (role, content). What this means is what you see in the code: instead of creating separate SystemMessagePromptTemplate and HumanMessagePromptTemplate objects, you just write (“system”, “…”) and (“human”, “…”) which provide a short-hand for those objects.

The roles are typically “system” (instructions), “human” or “user” (the question), and “ai” (for previous responses in multi-turn conversations).

Understanding Roles

Let’s make sure we understand these roles because the context is crucial.

The “system” role represents meta-instructions for the model. This is not actual user input, but rather the rules that shape how it behaves. Think of system messages as the model’s job description: they set its persona, expertise level, and constraints. These instructions are given higher priority by the model, being treated as foundational rules rather than conversation. The “human” or “user” messages, by contrast, are the actual questions from the person using the model. This is exactly equivalent to what someone types into a chat interface.

Now, could a user just write everything in one message? Sure! A user could say “Act as a relativity expert for non-specialists and explain time dilation.” But separating it into system + human messages gives us three advantages. It architecturally separates concerns in that the persona/rules live in one place, queries in another. It signals importance to the model since system messages are processed as “configuration” rather than “conversation,” carrying more weight. And, finally, it enables reusable templates because the system instruction stays constant while human queries change.

In real applications, this separation reflects how control is distributed: developers set the system message (defining the AI’s behavior), while end users provide the human message (asking their questions). The user never sees or controls the system prompt because it’s the internal ruleset for how the model should behave.

Thus, when testing a system, considering what internal rules are in place is extremely important. They are the baseline conditions upon which tests will execute. Even how they are worded can matter. Imagine if I changed my statement to simply “You are an expert in relativity” without including the bit about non-specialists.

Next Steps!

Here we did a dive into one part of the LangChain ecosystem: templates. This focus started to take us into another, which are the messages. In the next post, I’ll dig a bit more into the idea behind messages.

Share

This article was written by Jeff Nyman

Anything I put here is an approximation of the truth. You're getting a particular view of myself ... and it's the view I'm choosing to present to you. If you've never met me before in person, please realize I'm not the same in person as I am in writing. That's because I can only put part of myself down into words. If you have met me before in person then I'd ask you to consider that the view you've formed that way and the view you come to by reading what I say here may, in fact, both be true. I'd advise that you not automatically discard either viewpoint when they conflict or accept either as truth when they agree.

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.