Tool Open source
Apache Spark is a unified analytics engine for large-scale data processing, developed as an Apache project. It provides high-level APIs in Scala, Java, Python, and R, and executes general computation graphs for data analysis. Its higher-level components include Spark SQL and DataFrames, the pandas API on Spark, MLlib for machine learning, GraphX for graph processing, and Structured Streaming for stream processing. Spark can run locally or submit workloads to clusters through Spark, YARN, or Kubernetes integrations, and is used in large-scale data pipelines such as those underlying Poolside's model factory.
1 use taken from transcripts — each links to the moment in the video.
Supports the large-scale data pipelines used underneath the model factory.
1 in the library.