apache-superset - 技术专题

相关标签
reactpythonflaskdata-sciencebianalyticssupersetapachedata-visualizationdata-engineering

Here are 129 public repositories matching this topic...

An End-to-End ETL data pipeline that leverages pyspark parallel processing to process about 25 million rows of data coming from a SaaS application using Apache Airflow as an orchestration tool and various data warehouse technologies and finally using Apache Superset to connect to DWH for generating BI dashboards for weekly reports

  • Updated Dec 7, 2022
  • Python
aws-data-pipeline

A batch processing data pipeline, using AWS resources (S3, EMR, Redshift, EC2, IAM), provisioned via Terraform, and orchestrated from locally hosted Airflow containers. The end product is a Superset dashboard and a Postgres database, hosted on an EC2 instance at this address (powered down):

  • Updated May 14, 2022
  • Python

A Smart Traffic Management System for Ho Chi Minh City, Vietnam leveraging batch and real-time data processing, intuitive dashboards, and monitoring tools to optimize traffic flow, enhance safety, and support sustainable urban mobility through advanced analytics and user-friendly applications.

  • Updated Jan 17, 2025
  • Python

Production-ready Apache Superset with DuckLake integration. Stateless analytics architecture using DuckDB for compute, PostgreSQL for metadata, and S3/GCS/MinIO for data lake storage. Includes Docker Compose, Kubernetes Helm charts, BigQuery Integration, and CI/CD workflows. Supports MotherDuck cloud integration.

  • Updated Feb 6, 2026
  • Python