Langfuse v4: up to 165× faster · Read more
HandbookArchitecture

Langfuse Platform Architecture - High-Level Overview

Langfuse's infrastructure continuously evolves to support increasing scale and new product features. We started on Vercel and Supabase with Next.js and Postgres, and have evolved to the distributed architecture described below. As our product and scale requirements grow, we'll continue to mature our infrastructure to meet those needs.

Langfuse only depends on open source components and can be deployed locally, on cloud infrastructure, or on-premises.

Deployment architectureScroll to explore →
UI, API, and SDK clients connect to Langfuse Web. Web connects to PostgreSQL, Redis or Valkey, ClickHouse, and S3 or blob storage. Redis queues jobs for Langfuse Worker, which connects to PostgreSQL, ClickHouse, and S3. Both containers can optionally connect to an LLM API or gateway: Web for the playground, Worker for evaluations. The application containers and storage run in your VPC or on-premises. The optional LLM can also run in the same VPC or a peered VPC.YOUR INFRASTRUCTUREVPC / on-premisesUI, API & SDKsRequestsBACKGROUND PROCESSINGLangfuse WebUI & API serverlangfuse/langfuseLangfuse WorkerAsync event processinglangfuse/workerPostgreSQLTransactional dataRedis / ValkeyCache & job queueClickHouseObservability dataS3 / Blob storageEvents & attachmentsPlaygroundEvaluationsLLM API / GatewayOptional · bring your own
Web connectionsWorker connectionsOptional

Each component links to its setup guide. The LLM API can also run in the same VPC or a peered VPC.

Langfuse consists of two application containers, storage components, and an optional LLM API/Gateway.

  • Application Containers
    • Langfuse Web: The main web application serving the Langfuse UI and APIs.
    • Langfuse Worker: A worker that asynchronously processes events.
  • Storage Components:
    • Postgres: The main database for transactional workloads.
    • Clickhouse: High-performance OLAP database which stores traces, observations, and scores.
    • Redis/Valkey cache: A fast in-memory data structure store. Used for queue and cache operations.
    • S3/Blob Store: Object storage to persist all incoming events, multi-modal inputs, and large exports.
  • LLM API / Gateway: Some features depend on an external LLM API or gateway.

Langfuse can be deployed within a VPC or on-premises in high-security environments. Internet access is optional. See networking documentation for more details.

Infrastructure Components

Application Layer

  • Web container (NextJs): Serves the UI application and all APIs.
  • Worker container (Express): Processes ingestion events in the background and executes async tasks (e.g. exports, eval execution).

Evaluation Layer

  • LLM-as-a-Judge: The worker calls external LLM providers to execute model-based evaluators.
  • Code evaluator Lambda runners: Execute code evaluators outside the worker through the CodeEvalDispatcher abstraction. Production deployments use tenant-isolated AWS Lambda runners with AWS Lambda tenant isolation.

Storage Layer

  • PostgreSQL: Stores transactional data (users, organizations, projects, API keys, prompts, datasets, LLM as a Judge settings).
  • ClickHouse: Stores tracing data (traces, observations, scores). We use it to run dashboards, metrics, and render tables in the UI.
  • Redis: Stores event queue (BullMQ) and caching layer (API keys, prompts).
  • S3: Stores raw ingestion events and multi-modal attachments (images, audio).

Why do we need an OLAP database (Clickhouse) for observability data?

  • We built Langfuse initially on Postgres and eventually migrated to Clickhouse. We always knew that Postgres won't be the best fit for our observability data.
  • OLAP databases have a columnar layout. With that the database only scans data required to produce results for analytical queries (e.g. LLM cost over time).
  • We needed a multi-node database to scale our data insert.
  • As we are an open source product, we required a database which runs on an open source license.

Production environments

Our production infrastructure is deployed across multiple AWS regions with a fully automated CI/CD pipeline. All infrastructure is managed using Terraform. Cloudflare WAF (Web Application Firewall) serves as a central proxy in front of AWS.

Data Ingestion from SDKs

Architecture

  • SDKs: SDKs instrument the applications of our users. We built our own Python/JS SDKs which use OpenTelemetry under the hood.
  • API: SDKs send data to our API, which uploads the data to S3 and queues it for processing by the worker.
  • Redis queue: Decouples ingestion from processing. We only pass S3 references through Redis.
  • Worker processing: Asynchronously processes ingestion events, enriches events, flushes to ClickHouse.
  • Dual database: ClickHouse for analytical queries, Postgres for transactional data

Was this page helpful?

Last updated on