Job and task notifications, system destinations, duration warnings and how retries affect which alerts are sent.
From Ultra Transcenders DP-750 by Tony Rough (coming November 2026)
Notifications tell people or systems that a run started, succeeded, failed, ran too long or fell behind. They can be set on the whole job and on individual tasks; editing job notifications needs CAN MANAGE or IS OWNER on the job.
| Event | Fires when |
|---|---|
| Start | A run starts |
| Success | A run completes successfully, including Succeeded with failures |
| Failure | A run ends in an unsuccessful state |
| Duration warning | A run exceeds the Warning threshold of the Run duration metric |
| Streaming backlog | The average backlog over 10 minutes exceeds a threshold (repeats at 30-minute intervals while high); streaming observability is in Public Preview |
| Maintenance start / complete | A continuous job’s maintenance window begins or ends (system destinations only, not email) |
Destinations are email addresses or system destinations: Slack, Microsoft Teams, PagerDuty and HTTP webhooks. A workspace admin creates system destinations in the admin settings (Edit system notifications); each job or task can use at most three system destinations per event type. Webhooks receive a JSON payload with an event_type such as jobs.on_failure or jobs.on_duration_warning_threshold_exceeded. Because the content of Slack and Teams messages may change, build automation on a user-defined webhook instead.
Under Duration and streaming backlog thresholds (job) or Metric thresholds (task), the Run duration metric has two fields:
awaitTermination() in jobs.Common trap: Relying on a job-level failure notification to hear about every failed attempt - job-level notifications are not sent when failed tasks are retried. Add task-level notifications for per-attempt alerts, and use Mute notifications until the last retry if you want only the final outcome.
Common trap: Treating the Warning duration as a hard limit - it only raises a duration-warning event. Only Timeout stops the run and marks it Timed Out.
This note is one section of Ultra Transcenders DP-750: Implementing Data Engineering Solutions Using Azure Databricks, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Due on Amazon in November 2026, in Kindle and paperback editions.
About the book · DP-750 terms in the glossary · All DP-750 study notes
What standard (formerly shared) and dedicated (formerly single user) access modes allow, and when each is required.
Who manages the files, what DROP TABLE does to each, and why Databricks recommends managed tables.
How SQL UDF row filters and column masks restrict data per user, and how they differ from dynamic views.
How the two retention properties and VACUUM decide which table versions you can still query or restore.
SCD types 0, 1, 2 and others compared, and when to keep history in a dimension table.
The table-size thresholds for partitioning, partition sizing, and why liquid clustering is usually the better choice.
How expectations validate records in Lakeflow Spark Declarative Pipelines and what each violation action does.