Langfuse
Open source tracing, evaluation and prompt management for LLM apps


1/3Langfuse is an open source engineering platform for teams building LLM applications and agents, covering observability and evaluation over the same production data. Traces are hierarchical, so an agent run shows its LLM calls, tool invocations and retrieval steps in the shape they actually happened, and evaluation runs over them through LLM-as-a-judge, heuristic functions or human annotation queues. It also handles prompt management with versioning and one-click rollbacks, a playground for comparing models side by side on real inputs, experiments over test cases, and cost and latency dashboards with alerts. It is MIT licensed with over 100 integrations, and self-hosting is free.
- Hierarchical traces across LLM calls, tools and retrieval
- Evaluation by LLM-as-a-judge, heuristics or human review
- Prompt management with versions, deploys and rollbacks
- Playground compares models side by side on production inputs
- Experiments define test cases and compare the results
- Annotation queues for collaborative human review
- Cost and latency dashboards with automated alerts
- 100+ integrations, including LangChain and OpenTelemetry
What people are saying about Langfuse
72% neutral“* Langfuse is the only one with a real organization → project hierarchy; Laminar has workspaces with three roles.”
“I mean, for LLMs like via OpenRouter it's fairly easy (even Langfuse can estimate costs), but I also want to control costs on Apify (and it has a bunch of different actors with different cost per operations/volume/compute) - so , so far definitely no, but would really like to”
“📊 Visualize all agent activity in Datadog, Grafana, or Langfuse with OpenTelemetry.”
“Langfuse, Langsmith, Helicone, and just dumping tool calls to a Postgres table have all solved observability for agents already. Nobody serious was flying blind, you just were”
“real problem is real though, no clue which tool ate my budget on some dumb retry loop, but Helicone and Langfuse already do token tracking pretty well if that's what you need :)”
“and langfuse is open source and self hostable - did you consider running that yourself instead of rolling your own, or was your custom work already too far along by then?”
“Now it won't even respond to a Langfuse logfile chunk when my local system took down Tailscale. So even as a backup for when my local system has an issue it doesn't help.”
“We looked at LangSmith and LangFuse and several others. We used LangSmith for a while but we built a lot around it, but things went kind of ballistic on pricing which introduced too much unpredictability into the platform's operating costs, so we decided to roll our own.”
“Langfuse (30.2k★) AI observability with full prompt traces and evaluation scoring.”
“\- Use Observability tool such as Langsmith/LangFuse or MlFlow & measure whether your chatbot is actually working correctly or not?”
“For observability/tracing using langfuse or mlflow, you only add 1-2 line of codes and thats it.”
“@snr14 walks through building a stateful customer support agent with Python, LangGraph, and Langfuse.”
“Yes, there's a langfuse skill which you can use with Claude code, haven't set it up for codex but I'm sure it's something similar. Just note that you still need langfuse.”
“1)Observability which is done by langfuse but agentrealm tells in text what happens”
“I'm also experimenting with LangFuse for dev observability but haven't gotten much deeper than configuration yet.”
“Langfuse is insanely helpful if you want to improve your chat interface.”
“Is Hermes Cloud like a Digital Ocean VPS, so I can also install and host platforms like Hindsight, Matrix and Langfuse on it..?”
“langfuse 33.7k stars - self-hosted tracing, the layer where you read what your agent actually did. Part of ClickHouse since January, MIT except the ee/ folders”