Tool Open source

Apache Spark

Apache Spark is a unified analytics engine for large-scale data processing, developed as an Apache project. It provides high-level APIs in Scala, Java, Python, and R, and executes general computation graphs for data analysis. Its higher-level components include Spark SQL and DataFrames, the pandas API on Spark, MLlib for machine learning, GraphX for graph processing, and Structured Streaming for stream processing. Spark can run locally or submit workloads to clusters through Spark, YARN, or Kubernetes integrations, and is used in large-scale data pipelines such as those underlying Poolside's model factory.

View repository Visit site Mentioned in 1 video ↓

What Apache Spark is used for

1 use taken from transcripts — each links to the moment in the video.

  • Supports the large-scale data pipelines used underneath the model factory.

Videos mentioning Apache Spark

1 in the library.