OnCallReady

Lesson 25.4 · CI/CD Pipelines & Helm · 21 min read

Azure Pipelines YAML: triggers, stages, jobs, steps and the agent

In plain words

Think of a recipe book for a school kitchen. The book is split into chapters (stages): prepare, cook, serve. Each chapter has jobs, and every job is given to one cook at one table. The cook follows the steps in order at that table. Two cooks at different tables do not share a cutting board, so if the second needs the chopped onions, they must be put in a box and handed over.

In Azure Pipelines the book is azure-pipelines.yml, read from the commit that triggered the run. Stages hold jobs, a job runs on one agent (the azagent user on oncall-lab) with its own workspace, and steps are script, bash or tasks like Docker@2. The boxes are pipeline artifacts: publish in one job, download in the next.

The problem

Building and testing by hand on your laptop works until someone forgets, or their laptop differs from yours. A pipeline makes a server do it the same way on every push, and shows everyone the result. This lesson reads one pipeline file line by line.

What you need to know already: 25.1 (CI vs CD, artifact), 25.2 (branches, push), 6.1 (set -euo pipefail), 1.7 (exit codes).

Words for this lesson

The shape of a pipeline

azure-pipelines.yml lives in the app repo. Azure DevOps reads it from the commit that triggered the run - so a change to the pipeline is reviewed and versioned like code.

trigger:                      # CI trigger: which pushes start a run
  branches:
    include: [main]
  paths:
    exclude: ['*.md']         # a README change does not rebuild the image

pool: oncall-lab            # which agents may run it (a self-hosted pool)

stages:
- stage: Build                # stages run in order unless dependsOn says otherwise
  jobs:
  - job: build                # a job = one agent, one workspace, steps in sequence
    steps:
    - checkout: self          # implicit in a normal job; shown here for clarity
    - script: ./mvnw -B verify
      displayName: Build and test
    - publish: target
      artifact: jar

- stage: Package
  dependsOn: Build            # the default anyway: the previous stage
  jobs:
  - job: image
    steps:
    - download: current
      artifact: jar
    - bash: ls -l $(Pipeline.Workspace)/jar

Line by line:

The hierarchy is pipeline > stages > jobs > steps:

  stage   a boundary for approvals, environments and "what depends on what"
  job     runs on ONE agent; its steps share a workspace directory
  step    a script, a bash script, or a task (Docker@2, HelmDeploy@1, Cache@2 ...)

A workspace is the directory on the agent where a job's files live. Two jobs never share files, even in the same stage - a job gets a fresh workspace, maybe on a different machine. Files move between jobs as pipeline artifacts (publish / download). Forgetting that is the classic "it worked in the build job, the deploy job can't find target/app.jar".

Shortcuts: if the file has only steps: it is one job in one stage; only jobs: is one stage. Real pipelines use stages from the start because environments and approvals attach to stages.

The agent

Microsoft-hosted agents are fresh virtual machines Microsoft creates for each job and throws away after. In many companies, and in this lab, the agent is self-hosted: a service (a systemd unit, Ch 2) on a machine you control, registered to a pool. Here it runs on oncall-lab as the user azagent:

/home/azagent/myagent/              the agent (config.sh, run.sh, svc.sh)
/home/azagent/myagent/_work/1/      $(Pipeline.Workspace)
                         1/s/       $(Build.SourcesDirectory): the checkout
                         1/a/       $(Build.ArtifactStagingDirectory)
                         1/b/       $(Build.BinariesDirectory)
/home/azagent/myagent/_work/_temp/  scripts the agent generates for each step

config.sh registers the agent with Azure DevOps, run.sh runs it in the foreground, svc.sh installs it as a service. s = sources, a = artifact staging (where you put files to publish), b = binaries.

Self-hosted means state survives between runs (caches, Docker's layer cache, whatever the last job left in /tmp) and the agent's user has whatever rights you gave it. It is not you: azagent has no ~/.kube/config, no registry login, no git credentials - it gets those from service connections (lesson 25.7). A step that works in your terminal and fails in the pipeline is very often exactly that difference.

Predefined variables

Azure Pipelines sets these for every run; you read them as $(Name):

  $(Build.BuildId)            1041                    unique per run
  $(Build.BuildNumber)        20260924.3              name: controls the format
  $(Build.SourceBranch)       refs/heads/main
  $(Build.SourceBranchName)   main
  $(Build.SourceVersion)      9c41e7a2...             full commit sha
  $(Build.Reason)             IndividualCI | Manual | PullRequest
  $(Build.SourcesDirectory)   /home/azagent/myagent/_work/1/s
  $(Pipeline.Workspace)       /home/azagent/myagent/_work/1
  $(Agent.OS)                 Linux
  $(System.AccessToken)       (secret) the run's own token

refs/heads/main is git's full name for the branch main. Build.Reason says what started the run: a push (IndividualCI), a person (Manual) or a pull request. A token is a secret string that proves who you are to a server, like a password for a program.

Every non-secret variable is also an environment variable for scripts (Ch 6): dots become underscores, upper case: $(Build.SourceVersion) is $BUILD_SOURCEVERSION.

Steps: script, bash, task

- script: ./mvnw -B verify                 # CmdLine@2: bash on Linux
- bash: |                                  # Bash@3
    set -euo pipefail
    ./mvnw -B verify
    echo "built $(Build.SourceVersion)"
- task: Docker@2                           # a task: a packaged action with inputs
  inputs:
    containerRegistry: registry-lab
    repository: pay/payments-api
    command: buildAndPush
    tags: $(Build.BuildId)

The trap: a multi-line script: or bash: step is written to a file and run with bash - without set -e. The step's result is the exit code of the last command:

- script: |
    ./mvnw -B verify          # tests fail, exit 1 ...
    echo "build finished"     # ... but this exits 0, so the step is GREEN

Start every multi-line script with set -euo pipefail (Ch 6.1), or keep one command per step. GitLab CI behaves differently here - it fails on the first failing line - and people moving between the two get bitten both ways.

Reading the log

ci logs RUN JOB (simulator) prints the job log of run number RUN, job JOB, exactly as the agent writes it: an ISO timestamp with 7 decimals, ##[section] markers where each step starts and ends, the task banner, and ##[error] lines:

2026-09-24T10:00:05.8000466Z ##[section]Starting: Build and test
2026-09-24T10:00:05.8000466Z ==============================================================================
2026-09-24T10:00:05.8000466Z Task         : Command line
2026-09-24T10:00:05.8000466Z Description  : Run a command line script using Bash on Linux and macOS and cmd.exe on Windows
2026-09-24T10:00:05.8000466Z Version      : 2.250.1
...
2026-09-24T10:00:05.8000466Z Generating script.
2026-09-24T10:00:05.8000466Z ========================== Starting Command Output ===========================
2026-09-24T10:00:05.8000466Z /usr/bin/bash --noprofile --norc /home/azagent/myagent/_work/_temp/ad9c5b7a-aa33-4fd2-b9d2-824d82119421.sh
2026-09-24T10:00:34.8009520Z [ERROR] Tests run: 4, Failures: 1, Errors: 0, Skipped: 0
...
2026-09-24T10:00:34.8009520Z ##[error]Bash exited with code '1'.
2026-09-24T10:00:34.8009520Z ##[section]Finishing: Build and test

The /usr/bin/bash --noprofile --norc .../_temp/....sh line is the trap above made visible: your script was saved to a temp file and run with plain bash. Read logs from the first ##[error] upwards: the lines just above it say why (here: one test failed). ci show RUN gives the tree of stages, jobs and steps with their results - the view you get in the web UI.

Conditions and dependencies

A stage runs when everything in dependsOn succeeded - unless condition says otherwise:

- stage: Prod
  dependsOn: [Build, Dev]
  condition: and(succeeded(), eq(variables['Build.SourceBranch'], 'refs/heads/main'))

Read the condition as: "run only if the stages it depends on succeeded AND the branch is main". variables['X'] reads a variable inside a condition.

  succeeded()              all dependencies succeeded (the default)
  failed()                 at least one dependency failed
  always()                 run no matter what (clean-up, notifications)
  succeededOrFailed()      run unless the run was canceled
  eq/ne/and/or/not/in      string comparisons are case-INSENSITIVE
  startsWith, contains     startsWith(variables['Build.SourceBranch'], 'refs/heads/release/')

Once you write a custom condition, the default succeeded() is gone: a condition of just eq(variables['Build.SourceBranch'], 'refs/heads/main') runs the Prod stage even after Build failed. Always wrap it: and(succeeded(), ...).

Triggers

trigger: none                         # only manual (or scheduled) runs
trigger: [main, release/*]            # short form: branch list
trigger:
  branches:
    include: [main]
    exclude: [experimental/*]
  paths:
    include: [src, pom.xml, Dockerfile]

release/* matches every branch whose name starts with release/. The last form runs only for pushes to main that touched src, pom.xml (Maven's project file) or the Dockerfile.

A push to git.lab that matches starts a run with reason IndividualCI.

What you can now do

Why it helps

Most red pipelines you will debug come down to three things this lesson covers. First, "it works in the build job, the deploy job can't find target/app.jar": jobs never share files, so you need publish and download. Second, a green run with failing tests: a multi-line script: runs without set -e, so only the last command's exit code counts. Third, a Prod stage that ran after Build failed, because someone wrote a custom condition: and lost the implicit succeeded().

Knowing the agent's directory layout and that azagent has none of your credentials explains "works in my terminal, fails in the pipeline". And reading a log from the first ##[error] upwards is how you fix a teammate's broken build in minutes during a release freeze.

FAQ

What is the difference between a stage, a job and a step?

A stage is a boundary for ordering, approvals and environments; stages run in sequence by default. A job runs on one agent with one fresh workspace; jobs in a stage can run in parallel. A step is a single action inside a job: a script, a bash block or a task, run in order and sharing the job's files. Files cross jobs only as pipeline artifacts.

Why is my step green when the tests clearly failed?

A multi-line script: or bash: step is written to a temp file and run with bash --noprofile --norc, without set -e. The step's result is the exit code of the last command, so ./mvnw verify failing followed by echo done gives exit 0. Start every multi-line step with set -euo pipefail, or keep one command per step. GitLab behaves differently: it checks each line.

What is the difference between a Microsoft-hosted and a self-hosted agent?

A Microsoft-hosted agent is a fresh VM for every job, destroyed afterwards: clean but no warm caches and no access to your private network. A self-hosted agent is a service on a machine you manage, registered in a pool. State survives between runs (workspace, Docker layer cache, /tmp), it can reach internal systems, and its user has only the rights you gave it. Banks usually run self-hosted agents inside their network.

Why did my Prod stage run even though Build failed?

Because a custom condition: replaces the default succeeded() rather than adding to it. condition: eq(variables['Build.SourceBranch'], 'refs/heads/main') is true on main regardless of what failed before. Always write and(succeeded(), eq(...)). Use always() or failed() deliberately for clean-up and notification stages. String comparisons in these functions are case-insensitive.

How do I use a predefined variable like Build.SourceVersion inside a bash script?

Either as a macro, $(Build.SourceVersion), which Azure replaces in the script text before it runs, or as an environment variable: non-secret variables are exported with dots turned into underscores and the name upper-cased, so $BUILD_SOURCEVERSION. The environment form is safer in bash, because an unknown macro stays as literal $(Name) text, which bash then runs as command substitution.

In an interview Mid

Explain the structure of an Azure DevOps YAML pipeline.

azure-pipelines.yml lives in the app repo and is read from the commit that triggered the run:

The trap to mention: a multi-line script: runs without set -e, so its result is the last command's - tests fail, echo succeeds, the step is green. Start scripts with set -euo pipefail.

Also asked: A command works in your terminal but fails in the pipeline. What do you check? · Why can a pipeline step be green when the tests in it failed? · How do you stop a pipeline from deploying to production from a feature branch?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.