Speeds up SQL and DataFrame workloads on Azure Databricks using a native C++ engine that processes data in vectorised form. SQL warehouses and serverless use it automatically; through the Clusters or Jobs API, choose it by setting runtime_engine to PHOTON.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Photon in context, with comparison tables and the common traps.
Terms in this definition
- Serverless
Compute tier for single Azure SQL databases that scales automatically, pauses when idle and charges by the second. It is offered in General Purpose and Hyperscale, not Business Critical, and reserved capacity does not apply.
- DataFrame
Spark's rows-and-columns structure for data. In Structured Streaming it grows continually as each event adds rows, and queries run against it.
- Databricks
Analytics platform built on Apache Spark where data is processed in notebooks; Unity Catalog is the governance model it recommends.
- API
Short for application programming interface: a contract that client code calls programmatically, for example a web API secured with tokens or the Files, Images or Responses APIs.
Related terms
- SQL warehouse
Compute dedicated to running SQL queries, dashboards and BI tools, available as serverless, pro or classic. Serverless is the recommended type where offered, and every type runs Photon by default.
- vSphere Kubernetes releases
Ready-made node images, each a Kubernetes version and VKS core parts on Photon or Ubuntu, from which VKS clusters are built; shipped separately from vCenter.