
Langfuse: see what your AI chatbot does in production
Published on:
Reading time: 8 min
Topic: Technology
Author: Leandro Valencia
Your AI bot replies, but you cannot tell if it is right. What LLM observability is, and how open-source Langfuse shows cost, errors, and quality.
Table of Contents
- Three silent failures of an AI bot
- What LLM observability is
- When it is worth it
- Building is half the work
On a Monday the client writes: "The bot told a customer shipping was free. It is not." You open the panel and see the conversation. What you do not see is why it answered that.
Which prompt was live? Which documents did retrieval pass in? Was it the model, the context, or a change someone shipped on Friday? Without that, the fix is a guess.
This scene shows up again and again when you implement chatbots and WhatsApp flows for companies. Almost everything written about AI teaches you how to build the bot. Very little teaches you how to see what it is doing once real customers are talking to it. If you are still building that flow, the previous piece is how to turn a process into a chatbot. This is what comes after launch.
Three silent failures of an AI bot
A traditional bot fails out loud: it crashes, it returns an error, and someone notices. An AI bot fails quietly, because it always answers something. These are the three failures that show up most in production:
- Answers that sound right and are wrong. The model invents a price, a policy, or a schedule with full confidence. Nobody notices until a customer complains.
- Costs that jump with no visible cause. A conversation history that grows without a limit, or a loop inside an agent, doubles the API bill. The invoice arrives at the end of the month, when there is nothing left to cut.
- Prompts that get worse and nobody knows. You change the prompt to fix one case and break three others. Without comparable versions, users are the ones who tell you.
All three have the same cause. The problem is not only the model. It is that you cannot see what happened inside.
What LLM observability is
Think of an aircraft flight recorder. It does not prevent the accident. It lets you reconstruct, step by step, what happened. LLM observability does that for your bot: it stores every step of every conversation so you can review it later.
Each interaction is stored as a trace. The question "Is shipping free?" would look like this:
- Trace: the customer's WhatsApp message (1.8 s · USD 0.004)
- Knowledge-base search: the documents it found
- Model call: the exact prompt and its version, the reply, and the tokens
- Reply sent to the user
With that, the free-shipping case takes minutes. You see that search returned an expired promotion and the model treated it as current. You are no longer guessing. You know what to fix.
If you come from marketing, this is analytics for the conversations: who asked what, what it cost, and how well it went.
What Langfuse covers
Langfuse is an open-source platform, MIT licensed, for tracing, evaluating, and improving applications that use language models. Each feature answers one of the failures above.
| Feature | What it does | Which failure it addresses |
|---|---|---|
| Traces | Records every step of every conversation, with cost and latency | "I do not know why it answered that" |
| Prompt management | Versions prompts and lets you roll back to an earlier one | "The prompt got worse and I do not know when" |
| Evaluations | Scores replies with another model, with rules, or with a human review | "It sounds right but it is wrong" |
| Experiments | Tests a new prompt or model against real cases before you ship it | "I fixed one and broke three" |
| Dashboards | Cost, latency, and quality over time, with alerts | "The invoice doubled" |
It connects to OpenAI, Anthropic, Google, LangChain, the Vercel AI SDK, CrewAI, n8n, and any stack that speaks OpenTelemetry. The pricing page lists teams such as Canva, Khan Academy, and Hugging Face.
On January 16, 2026, ClickHouse acquired Langfuse. The project stays open source and self-hostable, under the same MIT license. In the announcement, ClickHouse put the project above 20,000 GitHub stars at the end of 2025. The acquisition does not change this guide: the cloud keeps the same endpoints, and the instrumentation code is the same against the cloud or against your own server.
What it costs
| Plan | Price | Included usage | Retention |
|---|---|---|---|
| Hobby | Free | 50,000 units/month, 2 users | 30 days |
| Core | USD 29/month | 100,000 units/month, unlimited users | 90 days |
| Pro | USD 199/month | 100,000 units/month, unlimited users | 3 years |
| Self-hosted | Free (you pay for the server) | No platform cap | Whatever you set |
Prices from the official page, checked in October 2026. Hobby has no extra units: when you hit the quota, the plan runs out. On Core and Pro the first 100,000 units are included and additional usage is billed in tiers, starting at USD 8 per 100,000. A unit is a tracing data point: the trace, each step, and each evaluation.
Hobby is enough to try it. Pick the region when you create the project. The cloud stores data in the United States, the European Union, or Japan, and the base URL changes with that choice.
Your first trace in ten minutes
The shortest path is a Python bot that already uses OpenAI. If you use n8n, LangChain, or another tool, the idea is the same and the detail is in the integrations guide.
- Create the account at cloud.langfuse.com, create a project, and copy both keys, the public one and the secret one.
- Install the package:
pip install langfuse openai
- Set the environment variables. The European Union URL is the one below. If you picked another region, use
https://us.cloud.langfuse.comorhttps://jp.cloud.langfuse.com.
LANGFUSE_PUBLIC_KEY="pk-lf-..."
LANGFUSE_SECRET_KEY="sk-lf-..."
LANGFUSE_BASE_URL="https://cloud.langfuse.com"
OPENAI_API_KEY="sk-..."
- Change the import and group the conversation. Langfuse ships a drop-in replacement for the OpenAI client that records the call.
propagate_attributesmarks the user and the session so the UI can group the turns of one chat. In a script that exits immediately,flush()pushes the data before the process dies.
from langfuse import observe, propagate_attributes, get_client
from langfuse.openai import openai # antes: import openai
@observe()
def responder(mensaje: str, usuario_id: str, conversacion_id: str):
with propagate_attributes(user_id=usuario_id, session_id=conversacion_id):
respuesta = openai.chat.completions.create(
name="respuesta-bot",
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "Eres el asistente de una tienda online."},
{"role": "user", "content": mensaje},
],
)
return respuesta.choices[0].message.content
print(responder("¿El envío es gratis?", "cliente-123", "chat-456"))
get_client().flush()
- Open the panel. In Tracing you will see the trace with the prompt, the reply, the tokens, the cost, and the time. With the
session_idyou can follow that user's full conversation.
If you want the data on your own server
You can self-host it with Docker. Locally, clone the repo, review the secrets marked CHANGEME in docker-compose.yml, and start the containers:
git clone https://github.com/langfuse/langfuse.git
cd langfuse
docker compose up
The docs say the web container is ready at http://localhost:3000 in about 2 or 3 minutes. There is no built-in admin user: the first visit is a sign-up. For a production machine, the same guide recommends at least 4 CPU cores and 16 GB of RAM, and it reminds you that Compose is not high availability. For a small company just starting, the free cloud plan is usually cheaper than that server.
When it is worth it
Langfuse is worth it from the moment the bot talks to real customers. Before that, you almost certainly do not need it.
Use it if:
- The bot serves customers or handles prices, orders, or support.
- You implement bots for other companies and you have to explain what happened when something fails.
- You spend more than a few dollars a month on AI APIs and you want to know on what.
- You change prompts often, or more than one person edits them.
You do not need it yet if:
- You are on a prototype for yourself or a weekend test.
- The bot platform is closed and will not let you connect an external tool. In that case, look at whatever logs the platform itself offers.
Alternatives. LangSmith and Braintrust cover similar ground. The practical difference with Langfuse is that it is open source and you can take the data to your own server.
Building is half the work
Shipping an AI bot takes an afternoon. Knowing whether it is working well is what separates an experiment from a product a customer can trust.
Langfuse does not make the model smarter. It does something more useful: it lets you see what is happening, so you fix the prompt, the documents, and the cost with the evidence in front of you. The entry plan is free. After the first customer complaint, flying blind no longer has an excuse.
Step-by-step guide
Create the Langfuse project
Sign in at cloud.langfuse.com, create a project, and copy the public and secret keys. Pick the region where the data will live.
Install the package
In the bot environment run pip install langfuse openai. Langfuse replaces the OpenAI client and records every call.
Set the environment variables
Define LANGFUSE_PUBLIC_KEY, LANGFUSE_SECRET_KEY, LANGFUSE_BASE_URL, and OPENAI_API_KEY before the process starts.
Wrap the model call
Import openai from langfuse.openai and propagate user_id and session_id so each conversation stays grouped.
Open the trace in the UI
Open Tracing, find the call, and read the prompt, the reply, the tokens, the cost, and the latency. In a short script, call flush before exit.
Frequently asked questions
What is Langfuse?
Langfuse is an open-source platform, MIT licensed, for tracing, evaluating, and improving applications that use language models. It records every step of a conversation, versions prompts, and shows cost, latency, and quality.
Is Langfuse free?
The Hobby plan costs nothing and does not ask for a card. It includes 50,000 units a month, 2 users, and 30 days of retention. You can also self-host it: the software is free and you pay for the server.
Does it work if my bot is not in Python?
Yes. This guide uses the Python SDK because it is the shortest path, but Langfuse also has a JavaScript SDK and integrations for LangChain, the Vercel AI SDK, CrewAI, n8n, and OpenTelemetry.
How is this different from the chatbot transcript?
The transcript shows the user message and the reply. A trace also shows which prompt was active, which documents retrieval returned, which model answered, how many tokens it spent, and how long it took.
Do I have to self-host it?
No. For a trial, and for most small companies, the free cloud plan is enough. Self-hosting makes sense when volume, data residency, or the cloud bill justifies a server.
Related Posts
Keep exploring similar content that may interest you

Opus 5.5 vs Sonnet 5.5: which Claude model to use in 2026
We compare Anthropic's models: Opus 5.5, Sonnet 5.5, Fable 5.1 and Haiku 4.5. Their prices, strengths and which one is worth using depending on the task.

The Witcher 3 Remastered: a lesson for your software
The Witcher 3 Remastered arrives for free with path tracing and rebuilt combat. Here is how the remaster model gives a second life to any software product.

Spec-Driven Development: a BDD, DDD, MDE & API-First guide
What Spec-Driven Development is, why it's back with AI, and how to use it with BDD, DDD, MDE and API-First, with examples, comparison table and agent workflow.
Partnerships
Tools I use every day, on better terms for this community.