Featured image for Langfuse: see what your AI chatbot does in production

Langfuse: see what your AI chatbot does in production

Published on:

Reading time: 8 min

Topic: Technology

Author: Leandro Valencia

#langfuse#llm observability#ai chatbot#prompt monitoring#openai costs

Your AI bot replies, but you cannot tell if it is right. What LLM observability is, and how open-source Langfuse shows cost, errors, and quality.

Table of Contents

On a Monday the client writes: "The bot told a customer shipping was free. It is not." You open the panel and see the conversation. What you do not see is why it answered that.

Which prompt was live? Which documents did retrieval pass in? Was it the model, the context, or a change someone shipped on Friday? Without that, the fix is a guess.

This scene shows up again and again when you implement chatbots and WhatsApp flows for companies. Almost everything written about AI teaches you how to build the bot. Very little teaches you how to see what it is doing once real customers are talking to it. If you are still building that flow, the previous piece is how to turn a process into a chatbot. This is what comes after launch.

Three silent failures of an AI bot

A traditional bot fails out loud: it crashes, it returns an error, and someone notices. An AI bot fails quietly, because it always answers something. These are the three failures that show up most in production:

  1. Answers that sound right and are wrong. The model invents a price, a policy, or a schedule with full confidence. Nobody notices until a customer complains.
  2. Costs that jump with no visible cause. A conversation history that grows without a limit, or a loop inside an agent, doubles the API bill. The invoice arrives at the end of the month, when there is nothing left to cut.
  3. Prompts that get worse and nobody knows. You change the prompt to fix one case and break three others. Without comparable versions, users are the ones who tell you.

All three have the same cause. The problem is not only the model. It is that you cannot see what happened inside.

What LLM observability is

Think of an aircraft flight recorder. It does not prevent the accident. It lets you reconstruct, step by step, what happened. LLM observability does that for your bot: it stores every step of every conversation so you can review it later.

Each interaction is stored as a trace. The question "Is shipping free?" would look like this:

  • Trace: the customer's WhatsApp message (1.8 s · USD 0.004)
    • Knowledge-base search: the documents it found
    • Model call: the exact prompt and its version, the reply, and the tokens
    • Reply sent to the user

With that, the free-shipping case takes minutes. You see that search returned an expired promotion and the model treated it as current. You are no longer guessing. You know what to fix.

If you come from marketing, this is analytics for the conversations: who asked what, what it cost, and how well it went.

What Langfuse covers

Langfuse is an open-source platform, MIT licensed, for tracing, evaluating, and improving applications that use language models. Each feature answers one of the failures above.

Feature What it does Which failure it addresses
Traces Records every step of every conversation, with cost and latency "I do not know why it answered that"
Prompt management Versions prompts and lets you roll back to an earlier one "The prompt got worse and I do not know when"
Evaluations Scores replies with another model, with rules, or with a human review "It sounds right but it is wrong"
Experiments Tests a new prompt or model against real cases before you ship it "I fixed one and broke three"
Dashboards Cost, latency, and quality over time, with alerts "The invoice doubled"

It connects to OpenAI, Anthropic, Google, LangChain, the Vercel AI SDK, CrewAI, n8n, and any stack that speaks OpenTelemetry. The pricing page lists teams such as Canva, Khan Academy, and Hugging Face.

On January 16, 2026, ClickHouse acquired Langfuse. The project stays open source and self-hostable, under the same MIT license. In the announcement, ClickHouse put the project above 20,000 GitHub stars at the end of 2025. The acquisition does not change this guide: the cloud keeps the same endpoints, and the instrumentation code is the same against the cloud or against your own server.

What it costs

Plan Price Included usage Retention
Hobby Free 50,000 units/month, 2 users 30 days
Core USD 29/month 100,000 units/month, unlimited users 90 days
Pro USD 199/month 100,000 units/month, unlimited users 3 years
Self-hosted Free (you pay for the server) No platform cap Whatever you set

Prices from the official page, checked in October 2026. Hobby has no extra units: when you hit the quota, the plan runs out. On Core and Pro the first 100,000 units are included and additional usage is billed in tiers, starting at USD 8 per 100,000. A unit is a tracing data point: the trace, each step, and each evaluation.

Hobby is enough to try it. Pick the region when you create the project. The cloud stores data in the United States, the European Union, or Japan, and the base URL changes with that choice.

Your first trace in ten minutes

The shortest path is a Python bot that already uses OpenAI. If you use n8n, LangChain, or another tool, the idea is the same and the detail is in the integrations guide.

  1. Create the account at cloud.langfuse.com, create a project, and copy both keys, the public one and the secret one.
  2. Install the package:
pip install langfuse openai
  1. Set the environment variables. The European Union URL is the one below. If you picked another region, use https://us.cloud.langfuse.com or https://jp.cloud.langfuse.com.
LANGFUSE_PUBLIC_KEY="pk-lf-..."
LANGFUSE_SECRET_KEY="sk-lf-..."
LANGFUSE_BASE_URL="https://cloud.langfuse.com"
OPENAI_API_KEY="sk-..."
  1. Change the import and group the conversation. Langfuse ships a drop-in replacement for the OpenAI client that records the call. propagate_attributes marks the user and the session so the UI can group the turns of one chat. In a script that exits immediately, flush() pushes the data before the process dies.
from langfuse import observe, propagate_attributes, get_client
from langfuse.openai import openai  # antes: import openai

@observe()
def responder(mensaje: str, usuario_id: str, conversacion_id: str):
    with propagate_attributes(user_id=usuario_id, session_id=conversacion_id):
        respuesta = openai.chat.completions.create(
            name="respuesta-bot",
            model="gpt-4o-mini",
            messages=[
                {"role": "system", "content": "Eres el asistente de una tienda online."},
                {"role": "user", "content": mensaje},
            ],
        )
        return respuesta.choices[0].message.content

print(responder("¿El envío es gratis?", "cliente-123", "chat-456"))
get_client().flush()
  1. Open the panel. In Tracing you will see the trace with the prompt, the reply, the tokens, the cost, and the time. With the session_id you can follow that user's full conversation.

If you want the data on your own server

You can self-host it with Docker. Locally, clone the repo, review the secrets marked CHANGEME in docker-compose.yml, and start the containers:

git clone https://github.com/langfuse/langfuse.git
cd langfuse
docker compose up

The docs say the web container is ready at http://localhost:3000 in about 2 or 3 minutes. There is no built-in admin user: the first visit is a sign-up. For a production machine, the same guide recommends at least 4 CPU cores and 16 GB of RAM, and it reminds you that Compose is not high availability. For a small company just starting, the free cloud plan is usually cheaper than that server.

When it is worth it

Langfuse is worth it from the moment the bot talks to real customers. Before that, you almost certainly do not need it.

Use it if:

  • The bot serves customers or handles prices, orders, or support.
  • You implement bots for other companies and you have to explain what happened when something fails.
  • You spend more than a few dollars a month on AI APIs and you want to know on what.
  • You change prompts often, or more than one person edits them.

You do not need it yet if:

  • You are on a prototype for yourself or a weekend test.
  • The bot platform is closed and will not let you connect an external tool. In that case, look at whatever logs the platform itself offers.

Alternatives. LangSmith and Braintrust cover similar ground. The practical difference with Langfuse is that it is open source and you can take the data to your own server.

Building is half the work

Shipping an AI bot takes an afternoon. Knowing whether it is working well is what separates an experiment from a product a customer can trust.

Langfuse does not make the model smarter. It does something more useful: it lets you see what is happening, so you fix the prompt, the documents, and the cost with the evidence in front of you. The entry plan is free. After the first customer complaint, flying blind no longer has an excuse.

Step-by-step guide

  1. Create the Langfuse project

    Sign in at cloud.langfuse.com, create a project, and copy the public and secret keys. Pick the region where the data will live.

  2. Install the package

    In the bot environment run pip install langfuse openai. Langfuse replaces the OpenAI client and records every call.

  3. Set the environment variables

    Define LANGFUSE_PUBLIC_KEY, LANGFUSE_SECRET_KEY, LANGFUSE_BASE_URL, and OPENAI_API_KEY before the process starts.

  4. Wrap the model call

    Import openai from langfuse.openai and propagate user_id and session_id so each conversation stays grouped.

  5. Open the trace in the UI

    Open Tracing, find the call, and read the prompt, the reply, the tokens, the cost, and the latency. In a short script, call flush before exit.

Frequently asked questions

What is Langfuse?

Langfuse is an open-source platform, MIT licensed, for tracing, evaluating, and improving applications that use language models. It records every step of a conversation, versions prompts, and shows cost, latency, and quality.

Is Langfuse free?

The Hobby plan costs nothing and does not ask for a card. It includes 50,000 units a month, 2 users, and 30 days of retention. You can also self-host it: the software is free and you pay for the server.

Does it work if my bot is not in Python?

Yes. This guide uses the Python SDK because it is the shortest path, but Langfuse also has a JavaScript SDK and integrations for LangChain, the Vercel AI SDK, CrewAI, n8n, and OpenTelemetry.

How is this different from the chatbot transcript?

The transcript shows the user message and the reply. A trace also shows which prompt was active, which documents retrieval returned, which model answered, how many tokens it spent, and how long it took.

Do I have to self-host it?

No. For a trial, and for most small companies, the free cloud plan is enough. Self-hosting makes sense when volume, data residency, or the cloud bill justifies a server.

Related Posts

Keep exploring similar content that may interest you

Partnerships

Tools I use every day, on better terms for this community.

Affiliate links. Your price does not change.See all partnerships
Training program

Ready to turn your idea into a real project?

Transforma is the program where you will learn to create, execute and scale your project with clarity and method.

Discover the Transforma Program