Spark's rows-and-columns structure for data. In Structured Streaming it grows continually as each event adds rows, and queries run against it.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains DataFrame in context, with comparison tables and the common traps.
Terms in this definition
- Structured Streaming
Spark's engine for near real-time data: you write a batch-like query and it processes new data incrementally, recording progress in a checkpoint with exactly-once guarantees. Auto Loader and streaming tables run on it.
- Event
Table in Log Analytics where entries from Windows event logs are kept.
Related terms
- Apache Spark cache
A cache you request explicitly, in memory or on disk, for a table or DataFrame using
df.cache(),df.persist()orCACHE TABLE, and release withunpersist. It differs from the automatic disk cache and cannot be used on serverless compute. - dropDuplicates
Removes repeated rows from a PySpark DataFrame, comparing every column or just the ones you pass, such as a business key. The row kept for each key is arbitrary.
- Intelligent cache
A Fabric Spark feature, enabled by default and allowed half of each node's disk, that keeps copies of OneLake files on the node's local SSD and swaps out stale ones automatically. It differs from calling cache() on a DataFrame, which is something you control in code.
- mltable Python package
Python library for working with MLTable blueprints. You call
mltable.load()on the folder holding the file, thento_pandas_dataframe()to materialise,take()to sample rows, orsave()to write a blueprint out. - Photon
Speeds up SQL and DataFrame workloads on Azure Databricks using a native C++ engine that processes data in vectorised form. SQL warehouses and serverless use it automatically; through the Clusters or Jobs API, choose it by setting
runtime_enginetoPHOTON.