11 Commits
Author SHA1 Message Date
Gitea Actions 3e25cf0b48 Auto-update blog content from Obsidian: 2026-08-21 20:36:36
Blog Deployment / Check-Rebuild (push) Successful in 5s
Blog Deployment / Build (push) Has been skipped
Blog Deployment / Deploy-Staging (push) Successful in 34s
Blog Deployment / Test-Staging (push) Successful in 1s
Blog Deployment / Merge (push) Successful in 5s
Blog Deployment / Deploy-Production (push) Successful in 34s
Blog Deployment / Test-Production (push) Successful in 2s
Blog Deployment / Clean (push) Has been skipped
Blog Deployment / Notify (push) Successful in 2s
2026-08-21 20:36:36 +00:00
Gitea Actions 79391d97ef Auto-update blog content from Obsidian: 2026-08-21 20:31:22
Blog Deployment / Check-Rebuild (push) Successful in 5s
Blog Deployment / Build (push) Successful in 25s
Blog Deployment / Deploy-Staging (push) Successful in 35s
Blog Deployment / Test-Staging (push) Successful in 2s
Blog Deployment / Merge (push) Successful in 6s
Blog Deployment / Deploy-Production (push) Successful in 33s
Blog Deployment / Test-Production (push) Successful in 2s
Blog Deployment / Clean (push) Successful in 2s
Blog Deployment / Notify (push) Successful in 2s
2026-08-21 20:31:22 +00:00
Gitea Actions 091b3b8e29 Auto-update blog content from Obsidian: 2026-08-21 07:25:55
Blog Deployment / Check-Rebuild (push) Failing after 2s
Blog Deployment / Build (push) Has been skipped
Blog Deployment / Deploy-Staging (push) Has been skipped
Blog Deployment / Test-Staging (push) Has been skipped
Blog Deployment / Merge (push) Has been skipped
Blog Deployment / Deploy-Production (push) Has been skipped
Blog Deployment / Test-Production (push) Has been skipped
Blog Deployment / Clean (push) Has been skipped
Blog Deployment / Notify (push) Successful in 2s
2026-08-21 07:25:55 +00:00
Gitea Actions 64c29c71db Auto-update blog content from Obsidian: 2026-08-01 20:55:28
Blog Deployment / Check-Rebuild (push) Successful in 14s
Blog Deployment / Build (push) Has been skipped
Blog Deployment / Deploy-Staging (push) Successful in 46s
Blog Deployment / Test-Staging (push) Successful in 2s
Blog Deployment / Merge (push) Successful in 14s
Blog Deployment / Deploy-Production (push) Successful in 49s
Blog Deployment / Test-Production (push) Successful in 2s
Blog Deployment / Clean (push) Has been skipped
Blog Deployment / Notify (push) Successful in 2s
2026-08-01 20:55:28 +00:00
Gitea Actions d7523ba978 Auto-update blog content from Obsidian: 2026-07-19 21:21:41
Blog Deployment / Check-Rebuild (push) Successful in 13s
Blog Deployment / Build (push) Has been skipped
Blog Deployment / Deploy-Staging (push) Successful in 49s
Blog Deployment / Test-Staging (push) Successful in 2s
Blog Deployment / Merge (push) Successful in 14s
Blog Deployment / Deploy-Production (push) Successful in 47s
Blog Deployment / Test-Production (push) Successful in 2s
Blog Deployment / Clean (push) Has been skipped
Blog Deployment / Notify (push) Successful in 3s
2026-07-19 21:21:41 +00:00
Gitea Actions 400f4abf95 Auto-update blog content from Obsidian: 2026-07-18 22:26:17
Blog Deployment / Check-Rebuild (push) Successful in 11s
Blog Deployment / Build (push) Has been skipped
Blog Deployment / Deploy-Staging (push) Successful in 49s
Blog Deployment / Test-Staging (push) Successful in 2s
Blog Deployment / Merge (push) Successful in 15s
Blog Deployment / Deploy-Production (push) Successful in 49s
Blog Deployment / Test-Production (push) Successful in 2s
Blog Deployment / Clean (push) Has been skipped
Blog Deployment / Notify (push) Successful in 3s
2026-07-18 22:26:17 +00:00
Gitea Actions 3240e029fa Auto-update blog content from Obsidian: 2026-07-18 21:26:18
Blog Deployment / Check-Rebuild (push) Successful in 13s
Blog Deployment / Build (push) Has been skipped
Blog Deployment / Deploy-Staging (push) Successful in 35s
Blog Deployment / Test-Staging (push) Successful in 2s
Blog Deployment / Merge (push) Successful in 15s
Blog Deployment / Deploy-Production (push) Successful in 35s
Blog Deployment / Test-Production (push) Successful in 3s
Blog Deployment / Clean (push) Has been skipped
Blog Deployment / Notify (push) Successful in 2s
2026-07-18 21:26:18 +00:00
Gitea Actions 57768b598f Auto-update blog content from Obsidian: 2026-06-09 20:24:47
Blog Deployment / Check-Rebuild (push) Successful in 6s
Blog Deployment / Build (push) Has been skipped
Blog Deployment / Deploy-Staging (push) Successful in 35s
Blog Deployment / Test-Staging (push) Successful in 2s
Blog Deployment / Merge (push) Successful in 7s
Blog Deployment / Deploy-Production (push) Successful in 35s
Blog Deployment / Test-Production (push) Successful in 2s
Blog Deployment / Clean (push) Has been skipped
Blog Deployment / Notify (push) Successful in 2s
2026-06-09 20:24:47 +00:00
Gitea Actions 2fac832345 Auto-update blog content from Obsidian: 2026-06-09 20:17:47
Blog Deployment / Deploy-Staging (push) Successful in 35s
Blog Deployment / Test-Staging (push) Successful in 3s
Blog Deployment / Merge (push) Successful in 7s
Blog Deployment / Deploy-Production (push) Successful in 35s
Blog Deployment / Test-Production (push) Successful in 3s
Blog Deployment / Clean (push) Has been skipped
Blog Deployment / Notify (push) Successful in 2s
Blog Deployment / Check-Rebuild (push) Successful in 6s
Blog Deployment / Build (push) Has been skipped
2026-06-09 20:17:47 +00:00
Gitea Actions bd6934ee8e Auto-update blog content from Obsidian: 2026-06-09 19:39:07
Blog Deployment / Check-Rebuild (push) Successful in 6s
Blog Deployment / Build (push) Successful in 30s
Blog Deployment / Deploy-Staging (push) Successful in 34s
Blog Deployment / Test-Staging (push) Successful in 2s
Blog Deployment / Merge (push) Successful in 6s
Blog Deployment / Deploy-Production (push) Successful in 34s
Blog Deployment / Test-Production (push) Successful in 3s
Blog Deployment / Clean (push) Successful in 2s
Blog Deployment / Notify (push) Successful in 3s
2026-06-09 19:39:07 +00:00
Vezpi e41c25b64f feat: secure docker image removal
Blog Deployment / Check-Rebuild (push) Successful in 6s
Blog Deployment / Build (push) Has been skipped
Blog Deployment / Deploy-Staging (push) Successful in 35s
Blog Deployment / Test-Staging (push) Successful in 2s
Blog Deployment / Merge (push) Successful in 8s
Blog Deployment / Deploy-Production (push) Successful in 35s
Blog Deployment / Test-Production (push) Successful in 2s
Blog Deployment / Clean (push) Has been skipped
Blog Deployment / Notify (push) Successful in 2s
2026-05-25 13:47:56 +00:00
13 changed files with 1532 additions and 1 deletions
+4 -1
View File
@@ -194,7 +194,10 @@ jobs:
steps: steps:
- name: Remove Old Docker Image - name: Remove Old Docker Image
run: | run: |
docker image rm $(docker image ls ${DOCKER_IMAGE} 2> /dev/null | awk '$NF != "U" && NR>1 {print $2}') IMAGE_IDS=$(docker image ls "${DOCKER_IMAGE}" 2>/dev/null | awk '$NF != "U" && NR>1 {print $2}')
if [ -n "$IMAGE_IDS" ]; then
docker image rm $IMAGE_IDS
fi
Notify: Notify:
needs: [Check-Rebuild, Build, Deploy-Staging, Test-Staging, Merge, Deploy-Production, Test-Production, Clean] needs: [Check-Rebuild, Build, Deploy-Staging, Test-Staging, Merge, Deploy-Production, Test-Production, Clean]
Binary file not shown.

After

Width:  |  Height:  |  Size: 28 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 67 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 157 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 39 KiB

@@ -0,0 +1,311 @@
---
slug: automating-proxmox-update-ansible
title: Automatiser les mises à jour de Proxmox VE avec Ansible
description: Automatisez les mises à jour dun cluster Proxmox VE avec Ansible, Semaphore UI et Ntfy, incluant les vérifications Ceph, les redémarrages progressifs et les rapports.
date: 2026-06-09
draft: false
tags:
- proxmox
- ansible
- semaphore-ui
- ntfy
categories:
- homelab
---
## Intro
Dans mon homelab, les mises à jour font partie de ces choses faciles à repousser.
Pas parce quelles sont compliquées, mais parce quelles sont manuelles. Je dois me connecter au bon système, vérifier l’état, appliquer les mises à jour, redémarrer si nécessaire, vérifier que tout revient correctement, puis répéter le même processus pour le composant suivant.
Et comme cest manuel, je le garde généralement pour plus tard.
Quand Proxmox VE 9.1 est sorti, je voulais déjà mettre à jour mon cluster, mais pas manuellement. Puis Proxmox VE 9.2 est devenu disponible il y a quelques jours, et je navais toujours pas construit de processus propre autour de cela. C’était un bon déclencheur pour enfin commencer à automatiser les mises à jour des parties importantes de mon homelab.
Lobjectif plus large est de simplifier et dautomatiser le patching de plusieurs composants clés :
- Proxmox VE
- OPNsense
- TrueNAS
Jai décidé de commencer par Proxmox parce quil est central dans le lab, et parce quun workflow de mise à jour progressive est un bon candidat pour lautomatisation.
---
## Les Outils Utilisés
Le processus de mise à jour est construit autour de quelques composants que jutilise déjà dans le lab.
[Proxmox VE](https://www.proxmox.com/en/proxmox-virtual-environment/overview) est ma plateforme de virtualisation. Le cluster utilise aussi Ceph, donc avant de toucher à un nœud, je veux massurer que le cluster est en bonne santé et que Ceph remonte `HEALTH_OK`.
[Ansible](https://docs.ansible.com/) est utilisé pour décrire le workflow de mise à jour sous forme de playbook.
[Semaphore UI](https://semaphoreui.com/) est utilisé pour exécuter le playbook depuis une interface web et le planifier.
[Ntfy](https://ntfy.sh/) est utilisé pour les notifications. Si les mises à jour sont planifiées, jai besoin de savoir quand quelque chose se passe, surtout si le cluster nest pas prêt ou si une mise à jour échoue.
---
## Création dun Topic Ntfy Dédié
Avant de planifier quoi que ce soit, je voulais un canal de notification dédié au homelab.
Jai créé un topic `homelab` dans Ntfy et un utilisateur dédié nommé `semaphore` avec un accès en écriture seule à ce topic.
```bash
ntfy user add semaphore
ntfy access semaphore homelab wo
```
Lidée est que Semaphore a uniquement besoin de publier des messages. Il na pas besoin dun accès en lecture.
Jai aussi ajouté le topic sur mon téléphone mobile afin de pouvoir recevoir des notifications lorsque lautomatisation sexécute.
Dans Semaphore, jai créé un groupe de variables nommé `Ntfy Homelab` pour stocker les valeurs nécessaires aux playbooks :
- `ntfy_url`
- `ntfy_topic`
- `ntfy_user`
- `NTFY_PASSWORD`
Le mot de passe est stocké comme variable denvironnement dans longlet `Secrets`.
![Groupe de variables Semaphore utilisé pour stocker la configuration Ntfy des notifications du homelab](images/semaphore-ntfy-homelab-variables.png)
---
## Conception du Workflow de Mise à Jour Proxmox
Pour Proxmox, je ne voulais pas dun playbook qui exécute simplement `apt upgrade` sur tous les nœuds. À la place, il fait les actions suivantes :
- Vérifier la santé du cluster
- Arrêter et envoyer une notification Ntfy si le cluster nest pas prêt
- Pour chaque nœud, vérifier si des mises à jour sont disponibles, et si oui :
- Activer le mode maintenance
- Attendre que les LXC et les VM quittent le nœud
- Mettre à jour les paquets
- Désactiver le rééquilibrage Ceph
- Redémarrer le nœud
- Activer le rééquilibrage Ceph
- Désactiver le mode maintenance
- Attendre que Ceph soit en bonne santé
- Envoyer un rapport Ntfy final
Le playbook complet est disponible sur mon [dépôt Homelab](https://github.com/Vezpi/Homelab/blob/main/ansible/proxmox/update_proxmox.yml)
---
## Détails du Workflow
Avant de démarrer la mise à jour progressive, le playbook vérifie :
- Le quorum du cluster Proxmox
- La santé de Ceph
Si lune de ces vérifications échoue, le playbook sarrête et envoie une notification Ntfy au lieu dessayer de continuer.
```yaml
- name: Verify cluster quorum
ansible.builtin.command: pvecm status
register: quorum_status
changed_when: false
failed_when: quorum_status.stdout is not search('Quorate:\\s*Yes')
- name: Verify Ceph health
ansible.builtin.command: ceph health
register: ceph_health
changed_when: false
failed_when: "'HEALTH_OK' not in ceph_health.stdout"
```
Cest une partie importante de lautomatisation. Une mise à jour planifiée ne doit pas continuer aveuglément si le cluster nest pas dans un bon état.
Le playbook met à jour les nœuds Proxmox avec `serial: 1`.
Cela signifie quun seul nœud est traité à la fois, ce qui est exactement ce que je veux pour une mise à jour de cluster.
Pour chaque nœud, le playbook commence par rafraîchir les dépôts et vérifie si des mises à jour sont disponibles en utilisant le mode check dAnsible.
```yaml
- name: Refresh repositories
ansible.builtin.apt:
update_cache: true
- name: Check if updates are available
ansible.builtin.apt:
upgrade: dist
check_mode: true
register: apt_check
```
Si aucune mise à jour nest disponible pour un nœud, la partie lourde du workflow est ignorée.
Si des mises à jour sont disponibles, le playbook stocke la version actuelle de Proxmox, active le mode maintenance, attend que les invités quittent le nœud, applique les mises à jour, redémarre le nœud, puis attend que Ceph soit de nouveau en bonne santé.
```yaml
- name: Enable maintenance mode
ansible.builtin.command: >
ha-manager crm-command node-maintenance enable {{ inventory_hostname_short }}
```
Une fois le mode maintenance activé, le playbook attend quil ne reste plus aucun LXC en cours dexécution sur le nœud :
```yaml
- name: Wait for LXCs to leave node
ansible.builtin.shell: |
pct list | awk 'NR>1 && $2=="running" {count++} END {print count+0}'
register: lxc_count
changed_when: false
until: lxc_count.stdout | int == 0
retries: 60
delay: 15
```
Il fait la même chose pour les VM en cours dexécution :
```yaml
- name: Wait for VMs to leave node
ansible.builtin.shell: |
qm list | awk 'NR>1 && $3=="running" {count++} END {print count+0}'
register: vm_count
changed_when: false
until: vm_count.stdout | int == 0
retries: 60
delay: 15
```
Une fois que le nœud est vide, la mise à niveau des paquets peut sexécuter :
```yaml
- name: Update packages
ansible.builtin.apt:
upgrade: full
autoremove: true
autoclean: true
```
Avant de redémarrer, le playbook définit `noout` sur les OSD Ceph :
```yaml
- name: Disable Ceph rebalancing
ansible.builtin.command: ceph osd set noout
```
Puis le nœud est redémarré :
```yaml
- name: Reboot node
ansible.builtin.reboot:
reboot_timeout: 900
post_reboot_delay: 30
```
Après le redémarrage, le rééquilibrage Ceph est réactivé, le mode maintenance est désactivé, et le playbook attend que Ceph revienne à `HEALTH_OK`.
```yaml
- name: Enable Ceph rebalancing
ansible.builtin.command: ceph osd unset noout
- name: Disable maintenance mode
ansible.builtin.command: >
ha-manager crm-command node-maintenance disable {{ inventory_hostname_short }}
- name: Wait for Ceph to be healthy
ansible.builtin.command: ceph health
register: ceph_status
changed_when: false
until: "'HEALTH_OK' in ceph_status.stdout"
retries: 60
delay: 15
delegate_to: "{{ groups['nodes'][0] }}"
```
Le résultat est une mise à jour progressive contrôlée au lieu dune procédure manuelle nœud par nœud.
---
## Envoi dun Rapport de Mise à Jour
À la fin du workflow, le playbook envoie un rapport via Ntfy. Il détermine dabord si au moins un nœud a été mis à jour :
```yaml
- name: Determine if updates occurred
ansible.builtin.set_fact:
updates_performed: "{{ groups['nodes'] | map('extract', hostvars) | selectattr('update_report', 'defined') | list | length > 0 }}"
```
Ensuite, il envoie un message au topic `homelab`.
Si aucune mise à jour n’était disponible, la notification lindique et utilise une priorité plus basse.
Si des mises à jour ont été appliquées, la notification liste les nœuds mis à jour et affiche la version de Proxmox avant et après la mise à jour.
La logique de rapport est basée sur le fact `update_report` sauvegardé pendant la mise à jour du nœud :
```yaml
- name: Save update report
ansible.builtin.set_fact:
update_report:
old: "{{ pve_old_version.stdout }}"
new: "{{ pve_new_version.stdout }}"
```
Le corps de la notification construit ensuite un résumé à partir de tous les nœuds :
```yaml
body: |
{% set updated_nodes = [] %}
{% for node in groups['nodes'] %}
{% if hostvars[node].update_report is defined %}
{% set _ = updated_nodes.append(node) %}
{% endif %}
{% endfor %}
{% if not updates_performed %}
No updates available on the cluster.
{% else %}
The following nodes were updated:
{% for node in updated_nodes %}
{% if hostvars[node].update_report.old == hostvars[node].update_report.new %}
- {{ hostvars[node].inventory_hostname_short }}: version {{ hostvars[node].update_report.old }} (unchanged)
{% else %}
- {{ hostvars[node].inventory_hostname_short }}: version {{ hostvars[node].update_report.old }} → {{ hostvars[node].update_report.new }}
{% endif %}
{% endfor %}
{% endif %}
```
Cela rend la tâche planifiée beaucoup plus facile à considérer comme fiable. Je nai pas besoin douvrir Semaphore à chaque fois pour savoir ce qui sest passé.
---
## Exécuter le Playbook depuis Semaphore
Une fois le playbook prêt, je lai poussé dans le dépôt et jai configuré un modèle de tâche Semaphore pour lexécuter.
![Modèle de tâche Semaphore utilisé pour exécuter le playbook de mise à jour Proxmox](images/semaphore-playbook-update-proxmox-template.png)
À partir de là, je pouvais lancer le workflow et le regarder agir sur le cluster.
Pendant lexécution, le nœud cible entre en mode maintenance et les workloads en cours dexécution sont migrés hors de celui-ci.
![Nœud Proxmox en mode maintenance pendant que le playbook de mise à jour migre les workloads hors du nœud](images/proxmox-update-playbook-maintenance.png)
Cest à ce moment-là que lautomatisation devient vraiment utile. Le playbook napplique pas seulement les mises à jour. Il prend aussi en charge les étapes opérationnelles autour de la mise à jour.
---
## Planification de la Mise à Jour
Après avoir affiné le playbook et validé le workflow, jai créé une planification dans Semaphore.
Dans `Schedule`, jai cliqué sur `New Schedule`, sélectionné `Cron`, donné un nom, sélectionné une planification hebdomadaire, puis choisi le vendredi à 4h00 UTC.
![Planification hebdomadaire Semaphore pour le playbook de mise à jour Proxmox](images/semaphore-schedule-proxmox-update.png)
À ce stade, le processus de mise à jour de Proxmox nest plus quelque chose dont je dois me souvenir pour le faire manuellement.
Il sexécute selon une planification, vérifie l’état du cluster avant de faire quoi que ce soit, met à jour un nœud à la fois, et envoie une notification avec le résultat.
---
## Conclusion
Ce projet est parti dun problème simple : je ne mettais pas mon homelab à jour régulièrement parce que le processus était encore trop manuel.
Lautomatisation des mises à jour Proxmox était la première étape importante. La partie importante n’était pas seulement dexécuter les mises à niveau de paquets, mais de les entourer des vérifications et des étapes opérationnelles qui ont du sens pour un cluster Proxmox avec Ceph.
Semaphore me donne une façon propre dexécuter et de planifier le playbook. Ansible décrit le processus de manière répétable. Ntfy boucle la boucle en me disant ce qui sest passé.
Les prochaines étapes logiques sont de continuer avec la même approche pour les autres composants clés du lab : OPNsense et TrueNAS.
@@ -0,0 +1,311 @@
---
slug: automating-proxmox-update-ansible
title: Automating Proxmox VE Updates with Ansible
description: Automate Proxmox VE cluster updates with Ansible, Semaphore UI and Ntfy, including Ceph checks, rolling reboots and reports.
date: 2026-06-09
draft: false
tags:
- proxmox
- ansible
- semaphore-ui
- ntfy
categories:
- homelab
---
## Intro
In my homelab, updates are one of those things that are easy to postpone.
Not because they are complicated, but because they are manual. I need to connect to the right system, check the state, apply the updates, reboot if needed, verify that everything comes back correctly, and then repeat the same process for the next component.
And because it is manual, I usually keep it for later.
When Proxmox VE 9.1 came out, I already wanted to update my cluster, but not manually. Then Proxmox VE 9.2 became available few days ago, and I still had not built a clean process around it. That was a good trigger to finally start automating updates for the important parts of my homelab.
The larger goal is to simplify and automate patching for several key components:
- Proxmox VE
- OPNsense
- TrueNAS
I decided to start with Proxmox because it is central to the lab, and because a rolling update workflow is a good candidate for automation.
---
## The Tools Involved
The update process is built around a few components I already use in the lab.
[Proxmox VE](https://www.proxmox.com/en/proxmox-virtual-environment/overview) is my virtualization platform. The cluster also uses Ceph, so before touching a node, I want to make sure the cluster is healthy and that Ceph is reporting `HEALTH_OK`.
[Ansible](https://docs.ansible.com/) is used to describe the update workflow as a playbook.
[Semaphore UI](https://semaphoreui.com/) is used to run the playbook from a web interface and schedule it.
[Ntfy](https://ntfy.sh/) is used for notifications. If updates are scheduled, I need to know when something happens, especially if the cluster is not ready or if an update fails.
---
## Creating a Dedicated Ntfy Topic
Before scheduling anything, I wanted a notification channel dedicated to the homelab.
I created a `homelab` topic in Ntfy and a dedicated user named `semaphore` with write-only access to this topic.
```bash
ntfy user add semaphore
ntfy access semaphore homelab wo
```
The idea is that Semaphore only needs to publish messages. It does not need read access.
I also added the topic on my mobile phone so I can receive notifications when the automation runs.
In Semaphore, I created a variable group named `Ntfy Homelab` to store the values needed by the playbooks:
- `ntfy_url`
- `ntfy_topic`
- `ntfy_user`
- `NTFY_PASSWORD`
The password is stored as an environment variable in the `Secrets` tab.
![Semaphore variable group used to store the Ntfy configuration for homelab notifications](images/semaphore-ntfy-homelab-variables.png)
---
## Designing the Proxmox Update Workflow
For Proxmox, I did not want a playbook that simply runs `apt upgrade` on all nodes. Instead, it doing the following:
- Check cluster health
- Stop and send a Ntfy notification if the cluster is not ready
- For each node, check if updates are available, if so:
- Enable maintenance mode
- Wait for LXCs and VMs to leave the node
- Update packages
- Disable Ceph rebalancing
- Reboot the node
- Enable Ceph rebalancing
- Disable maintenance mode
- Wait for Ceph to be healthy
- Send a final Ntfy report
The full playbook is available on my [Homelab repo](https://github.com/Vezpi/Homelab/blob/main/ansible/proxmox/update_proxmox.yml)
---
## Workflow Details
Before starting the rolling update, the playbook checks:
- Proxmox cluster quorum
- Ceph health
If one of these checks fails, the playbook stops and sends a Ntfy notification instead of trying to continue.
```yaml
- name: Verify cluster quorum
ansible.builtin.command: pvecm status
register: quorum_status
changed_when: false
failed_when: quorum_status.stdout is not search('Quorate:\\s*Yes')
- name: Verify Ceph health
ansible.builtin.command: ceph health
register: ceph_health
changed_when: false
failed_when: "'HEALTH_OK' not in ceph_health.stdout"
```
This is an important part of the automation. A scheduled update should not blindly continue if the cluster is not in a good state.
The playbook updates the Proxmox nodes with `serial: 1`.
That means only one node is handled at a time, which is exactly what I want for a cluster update.
For each node, the playbook first refreshes the repositories and checks if updates are available using Ansible check mode.
```yaml
- name: Refresh repositories
ansible.builtin.apt:
update_cache: true
- name: Check if updates are available
ansible.builtin.apt:
upgrade: dist
check_mode: true
register: apt_check
```
If no updates are available for a node, the heavy part of the workflow is skipped.
If updates are available, the playbook stores the current Proxmox version, enables maintenance mode, waits for guests to leave the node, applies the updates, reboots the node, and then waits for Ceph to be healthy again.
```yaml
- name: Enable maintenance mode
ansible.builtin.command: >
ha-manager crm-command node-maintenance enable {{ inventory_hostname_short }}
```
After maintenance mode is enabled, the playbook waits until no running LXC remains on the node:
```yaml
- name: Wait for LXCs to leave node
ansible.builtin.shell: |
pct list | awk 'NR>1 && $2=="running" {count++} END {print count+0}'
register: lxc_count
changed_when: false
until: lxc_count.stdout | int == 0
retries: 60
delay: 15
```
It does the same for running VMs:
```yaml
- name: Wait for VMs to leave node
ansible.builtin.shell: |
qm list | awk 'NR>1 && $3=="running" {count++} END {print count+0}'
register: vm_count
changed_when: false
until: vm_count.stdout | int == 0
retries: 60
delay: 15
```
Once the node is empty, the package upgrade can run:
```yaml
- name: Update packages
ansible.builtin.apt:
upgrade: full
autoremove: true
autoclean: true
```
Before rebooting, the playbook sets Ceph OSD `noout`:
```yaml
- name: Disable Ceph rebalancing
ansible.builtin.command: ceph osd set noout
```
Then the node is rebooted:
```yaml
- name: Reboot node
ansible.builtin.reboot:
reboot_timeout: 900
post_reboot_delay: 30
```
After the reboot, Ceph rebalancing is enabled again, maintenance mode is disabled, and the playbook waits for Ceph to return to `HEALTH_OK`.
```yaml
- name: Enable Ceph rebalancing
ansible.builtin.command: ceph osd unset noout
- name: Disable maintenance mode
ansible.builtin.command: >
ha-manager crm-command node-maintenance disable {{ inventory_hostname_short }}
- name: Wait for Ceph to be healthy
ansible.builtin.command: ceph health
register: ceph_status
changed_when: false
until: "'HEALTH_OK' in ceph_status.stdout"
retries: 60
delay: 15
delegate_to: "{{ groups['nodes'][0] }}"
```
The result is a controlled rolling update instead of a manual node-by-node procedure.
---
## Sending an Update Report
At the end of the workflow, the playbook sends a report through Ntfy. It first determines if at least one node was updated:
```yaml
- name: Determine if updates occurred
ansible.builtin.set_fact:
updates_performed: "{{ groups['nodes'] | map('extract', hostvars) | selectattr('update_report', 'defined') | list | length > 0 }}"
```
Then it sends a message to the `homelab` topic.
If no updates were available, the notification says so and uses a lower priority.
If updates were applied, the notification lists the updated nodes and shows the Proxmox version before and after the update.
The report logic is based on the `update_report` fact saved during the node update:
```yaml
- name: Save update report
ansible.builtin.set_fact:
update_report:
old: "{{ pve_old_version.stdout }}"
new: "{{ pve_new_version.stdout }}"
```
The notification body then builds a summary from all nodes:
```yaml
body: |
{% set updated_nodes = [] %}
{% for node in groups['nodes'] %}
{% if hostvars[node].update_report is defined %}
{% set _ = updated_nodes.append(node) %}
{% endif %}
{% endfor %}
{% if not updates_performed %}
No updates available on the cluster.
{% else %}
The following nodes were updated:
{% for node in updated_nodes %}
{% if hostvars[node].update_report.old == hostvars[node].update_report.new %}
- {{ hostvars[node].inventory_hostname_short }}: version {{ hostvars[node].update_report.old }} (unchanged)
{% else %}
- {{ hostvars[node].inventory_hostname_short }}: version {{ hostvars[node].update_report.old }} → {{ hostvars[node].update_report.new }}
{% endif %}
{% endfor %}
{% endif %}
```
This makes the scheduled job much easier to trust. I do not need to open Semaphore every time to know what happened.
---
## Running the Playbook from Semaphore
Once the playbook was ready, I pushed it to the repository and configured a Semaphore task template to run it.
![Semaphore task template used to run the Proxmox update playbook](images/semaphore-playbook-update-proxmox-template.png)
From there, I could launch the workflow and watch it act on the cluster.
During execution, the target node enters maintenance mode and the running workloads are migrated away from it.
![Proxmox node in maintenance mode while the update playbook migrates workloads away](images/proxmox-update-playbook-maintenance.png)
This is the point where the automation becomes really useful. The playbook is not only applying updates. It is also taking care of the operational steps around the update.
---
## Scheduling the Update
After refining the playbook and validating the workflow, I created a schedule in Semaphore.
In `Schedule`, I clicked `New Schedule`, selected `Cron`, gave it a name, selected a weekly schedule, and picked Friday at 4:00 AM UTC.
![Weekly Semaphore schedule for the Proxmox update playbook](images/semaphore-schedule-proxmox-update.png)
At this point, the Proxmox update process is no longer something I need to remember to do manually.
It runs on a schedule, checks the state of the cluster before doing anything, updates one node at a time, and sends a notification with the result.
---
## Conclusion
This project started with a simple problem: I was not updating my homelab regularly because the process was still too manual.
Automating Proxmox updates was the first milestone. The important part was not only running package upgrades, but wrapping them in the checks and operational steps that make sense for a Proxmox cluster with Ceph.
Semaphore gives me a clean way to run and schedule the playbook. Ansible describes the process in a repeatable way. Ntfy closes the loop by telling me what happened.
The next logical steps are to continue the same approach for the other key components of the lab: OPNsense and TrueNAS.
Binary file not shown.

After

Width:  |  Height:  |  Size: 69 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 148 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 162 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 42 KiB

@@ -0,0 +1,453 @@
---
slug: automating-opnsense-update-ansible
title: Automatiser les mises à jour d'OPNsense avec Ansible
description: Automatiser les mises à jour d'un cluster HA OPNsense dans un homelab avec Ansible, Semaphore UI, des vérifications CARP, des snapshots Proxmox et des notifications Ntfy.
date: 2026-08-21
draft: false
tags:
- opnsense
- ansible
- semaphore-ui
- ntfy
- proxmox
categories:
- homelab
---
## Intro
Dans mon homelab, la plupart des composants de l'infrastructure sont déjà suffisamment redondants pour tolérer les opérations de maintenance, mais le processus de mise à jour lui-même restait encore trop manuel. Proxmox a été le premier élément que j'ai automatisé dans ce [billet]({{< ref "post/20-automating-proxmox-update-ansible" >}}). Une fois ce workflow exécuté par Semaphore selon un calendrier, OPNsense est devenu la prochaine cible logique.
Mon installation OPNsense est un cluster HA composé de deux nœuds. Le nœud maître, `cerbere-head1`, fonctionne sur Proxmox. Le nœud de secours, `cerbere-head2`, fonctionne sur TrueNAS. Cette séparation est volontaire, car je veux que le réseau puisse survivre à une maintenance ou à une panne du cluster Proxmox.
L'objectif est simple : créer un playbook Ansible capable de mettre à jour ou de faire évoluer le cluster HA OPNsense de manière sûre, dans le bon ordre, avec des vérifications avant toute modification et une notification à la fin.
---
## Stratégie de mise à jour
OPNsense expose une API qui permet de vérifier l'état du système, l'état du firmware, les services et l'état des adresses IP virtuelles CARP. Ansible pilote l'automatisation au moyen d'appels API, tandis que Semaphore UI sert de contrôleur pour exécuter le playbook manuellement ou selon un calendrier. Ntfy est utilisé pour le rapport final et les notifications d'échec.
Pour le nœud hébergé sur Proxmox, j'utilise également la [collection Ansible community.proxmox](https://docs.ansible.com/projects/ansible/latest/collections/community/proxmox/index.html). Elle permet au playbook de créer un snapshot de la VM avant de mettre à jour le nœud maître du firewall et de revenir à ce snapshot si nécessaire.
Le point important est que les deux nœuds OPNsense ne sont pas traités exactement de la même manière. Le nœud de secours fonctionne sur TrueNAS, le playbook le met donc à jour sans créer de snapshot d'hyperviseur. Le nœud maître fonctionne sur Proxmox, le playbook crée donc un snapshot avant de démarrer l'opération sur le firmware.
---
## Création d'un utilisateur API dans OPNsense
Pour permettre à Ansible d'interagir avec OPNsense, je crée un utilisateur dédié sur le nœud maître.
Cet utilisateur s'appelle `automation`, avec un mot de passe aléatoire et uniquement les privilèges nécessaires au playbook :
- `Interface: Virtual IPS: Status`
- `System: Firmware`
- `System: Status`
- `Status: Services`
Comme il s'agit d'un cluster HA, la synchronisation entre les deux nœuds OPNsense gère la création de l'utilisateur sur le nœud de secours.
Après avoir créé l'utilisateur, je génère une clé API.
![opnsense-user-create-api-key.png](images/opnsense-user-create-api-key.png)
La clé API OPNsense est générée depuis l'utilisateur dédié à l'automatisation.
Le fichier téléchargé contient la `key` et le `secret` de l'API. Dans l'interface OPNsense, la clé reste visible dans l'onglet `ApiKeys`, mais le secret ne l'est plus.
Je teste d'abord les appels API avec Bruno depuis VS Code. Une fois les appels de base fonctionnels, je charge les identifiants dans Semaphore.
---
## Préparation de Semaphore UI
Dans Semaphore, je crée une entrée dans le magasin de clés nommée `OPNsense automation`, qui contient la clé et le secret de l'API.
Je crée ensuite un inventaire pour les nœuds OPNsense :
```yaml
---
all:
children:
opnsense:
vars:
ansible_connection: local
children:
opnsense_backup:
hosts:
cerbere-head2:
ansible_host: 192.168.88.3
main_role: BACKUP
hypervisor: TrueNAS
opnsense_master:
hosts:
cerbere-head1:
ansible_host: 192.168.88.2
main_role: MASTER
hypervisor: Proxmox
proxmox_vmid: 122
```
Le playbook s'exécute localement depuis Semaphore et communique avec chaque firewall via l'API OPNsense.
Je crée également un groupe de variables nommé `OPNsense automation API`, avec les identifiants API et quelques variables partagées :
- `OPNSENSE_API_KEY`
- `OPNSENSE_API_SECRET`
- `opnsense_api_key`
- `opnsense_api_secret`
- `opnsense_https_port`
- `opnsense_host`
Le port HTTPS est configuré sur `4443`, et l'hôte est construit à partir de l'adresse présente dans l'inventaire et de ce port.
Enfin, je crée le modèle de tâche Semaphore.
![semaphore-new-template-task-opnsense-update.png](images/semaphore-new-template-task-opnsense-update.png)
Le modèle de tâche Semaphore utilisé pour exécuter le playbook de mise à jour OPNsense.
Avant d'aller plus loin, je vérifie qu'Ansible peut interroger les deux nœuds :
```yaml
- name: Check node availability
ansible.builtin.uri:
url: "https://{{ opnsense_host }}/api/core/system/status"
method: GET
user: "{{ opnsense_api_key }}"
password: "{{ opnsense_api_secret }}"
force_basic_auth: true
validate_certs: false
```
À ce stade, l'automatisation peut atteindre les deux nœuds et s'authentifier auprès de leur API.
---
## Rendre la maintenance CARP utilisable depuis l'API
La maintenance CARP est un élément important du workflow.
Avant de mettre à jour un nœud, je veux le placer en mode maintenance afin que les adresses IP virtuelles ne soient pas actives sur le nœud en cours de mise à jour. C'est particulièrement important lors de la mise à jour du nœud maître, car le nœud de secours doit prendre le relais proprement avant le début de l'opération.
Pendant les tests de l'API, l'endpoint de maintenance CARP renvoyait `403 Forbidden` :
```text
POST https://{{opnsense_host}}/api/diagnostics/interface/carp_status/maintenance
```
Les privilèges avaient bien été accordés dans l'interface Web, le problème semblait donc lié à la définition de l'ACL. J'ai modifié manuellement le fichier ACL d'OPNsense afin d'ajouter le wildcard au motif de l'API :
```xml
<pattern>api/diagnostics/interface/carp_status/*</pattern>
```
Après avoir redémarré le système, l'endpoint a renvoyé une réponse correcte.
J'ai créé une petite [PR](https://github.com/opnsense/core/pull/10428) dans le projet OPNsense pour corriger le problème. Elle a été rapidement fusionnée dans la branche `master` de `opnsense/core`. C'était ma première contribution à ce projet.
![github-opnsense-pr-merged.png](images/github-opnsense-pr-merged.png)
La petite correction de l'ACL OPNsense a été fusionnée en amont.
Plus tard, OPNsense [26.1.11](https://forum.opnsense.org/index.php?topic=52257.0) est sorti avec cette correction. J'ai alors pu tester le playbook complet sans dépendre de la modification manuelle de l'ACL.
---
## Conception du workflow du playbook
Le playbook suit un ordre simple :
- Vérifier l'état des deux nœuds
- Mettre d'abord à jour le nœud de secours
- Mettre ensuite à jour le nœud maître
- Envoyer une notification finale
Cet ordre est important. Le nœud de secours est mis à jour en premier, car le maître reste actif. Ensuite, avant de mettre à jour le maître, le playbook active le mode maintenance CARP afin de permettre au nœud de secours de prendre le relais.
Le playbook prend également en charge plusieurs actions au moyen d'un sondage Semaphore.
![semaphore-opnsense-update-survey-action.png](images/semaphore-opnsense-update-survey-action.png)
Le sondage Semaphore me permet de choisir entre une vérification, une mise à jour et une évolution de version.
C'est nécessaire, car OPNsense n'expose pas les mises à jour et les évolutions de version exactement de la même manière. Les variables correspondant à la version cible et au redémarrage requis diffèrent selon l'action choisie. Le playbook résout ces différences avant de déterminer quoi faire.
---
## Vérifications du firmware et de CARP
La première phase s'exécute sur les deux nœuds OPNsense.
Elle récupère des informations au moyen de plusieurs appels API :
- État du système
- Vérification du firmware
- État du firmware
- État des adresses IP virtuelles CARP
Le playbook enregistre le résultat sous forme de facts réutilisables plus tard :
```yaml
- name: Store node facts
ansible.builtin.set_fact:
firmware_action: "{{ opnsense_action | default('update')}}"
firmware_status: "{{ _firmware_status.json.status }}"
firmware_status_msg: "{{ _firmware_status.json.status_msg }}"
firmware_current_version: "{{ _firmware_status.json.product.product_version | default('unknown') }}"
firmware_product_series: "{{ _firmware_status.json.product.product_series | default('unknown') }}"
firmware_update_version: "{{ _firmware_status.json.upgrade_packages | selectattr('name', 'equalto', 'opnsense') | map(attribute='new_version') | first | default('') }}"
firmware_upgrade_version: "{{ _firmware_status.json.upgrade_major_version | default('unknown') }}"
firmware_upgrade_message: "{{ _firmware_status.json.upgrade_major_message | regex_replace('<[^>]+>', ' ') | regex_replace('\\s{2,}', ' ') | trim | default('') }}"
firmware_up_to_date: "{{ _firmware_status.json.status_msg == 'There are no updates available on the selected mirror.' }}"
needs_reboot: "{{ (_firmware_status.json.upgrade_needs_reboot == '1') if _firmware_status.json.status == 'upgrade' else (_firmware_status.json.needs_reboot == '1') }}"
total_vips: "{{ _vip_status.json.rowCount }}"
mismatched_vips: "{{ _vip_status.json.rows | rejectattr('status', 'equalto', main_role) | list | length }}"
in_maintenance: "{{ _vip_status.json.carp.maintenancemode }}"
```
Le playbook détermine ensuite s'il s'agit d'une version précise ou d'une série de versions :
```yaml
- name: Resolve update-vs-upgrade specifics
ansible.builtin.set_fact:
firmware_target_kind: "{{ 'series' if firmware_status == 'upgrade' else 'version' }}"
firmware_target_value: "{{ firmware_upgrade_version if firmware_status == 'upgrade' else firmware_update_version }}"
```
Cela simplifie la gestion du reste du playbook. Il peut ensuite vérifier que le nœud a atteint la `version` ou la `series` attendue sans dupliquer toute la logique.
La première phase valide également plusieurs conditions avant de continuer :
- Le nœud ne doit pas être déjà en mode maintenance CARP
- Les adresses IP virtuelles doivent correspondre au rôle attendu
- Au moins une adresse IP virtuelle doit être gérée
Si l'une de ces vérifications échoue, le playbook s'arrête et envoie une notification Ntfy.
## Gestion correcte des conditions de non-exécution
L'un des aspects les plus délicats n'est pas la mise à jour elle-même, mais la décision de ne pas mettre à jour.
Après une première exécution réussie, il est possible que le nœud de secours soit déjà à jour alors que le maître ne l'est pas encore. L'exécution suivante ne doit donc pas mettre à jour le nœud de secours une nouvelle fois s'il possède déjà la version ciblée par le maître.
J'ajoute une condition de non-exécution pour ce cas :
```yaml
- name: Backup already updated
ansible.builtin.set_fact:
skip_update: true
firmware_status: "skipped"
delegate_to: "{{ groups['opnsense_backup'][0] }}"
delegate_facts: true
run_once: true
when: >-
(hostvars[groups['opnsense_backup'][0]].firmware_current_version
if hostvars[groups['opnsense_master'][0]].firmware_target_kind == 'version'
else hostvars[groups['opnsense_backup'][0]].firmware_product_series)
== hostvars[groups['opnsense_master'][0]].firmware_target_value
```
Je généralise ensuite ce comportement.
Le playbook ignore un nœud dans les cas suivants :
- Aucune mise à jour n'est disponible
- Une évolution de version est disponible, mais l'action demandée est une mise à jour
- Une mise à jour est disponible, mais l'action demandée est une évolution de version
- Le nœud de secours possède déjà la version ou la série ciblée par le maître
La notification finale est ainsi beaucoup plus claire, car un nœud ignoré n'est pas considéré comme une erreur. Elle indique simplement qu'aucune action n'était nécessaire.
## Mise à jour du nœud de secours
Le nœud de secours fonctionne sur TrueNAS, cette phase ne crée donc pas de snapshot d'hyperviseur.
Le playbook active le mode maintenance CARP, déclenche l'opération sur le firmware, attend le début de la mise à jour, attend le redémarrage du nœud si nécessaire, puis attend que celui-ci soit de nouveau disponible.
La partie correspondante ressemble à ceci :
```yaml
- name: Trigger firmware {{ firmware_action }}
ansible.builtin.uri:
url: "https://{{ opnsense_host }}/api/core/firmware/{{ firmware_action }}"
method: POST
user: "{{ opnsense_api_key }}"
password: "{{ opnsense_api_secret }}"
force_basic_auth: true
validate_certs: false
```
Si un redémarrage est nécessaire, le playbook attend que le port HTTPS ne soit plus accessible :
```yaml
- name: Wait for node to reboot after the {{ firmware_action }}
ansible.builtin.wait_for:
host: "{{ ansible_host }}"
port: "{{ opnsense_https_port }}"
state: stopped
timeout: 3600
when: needs_reboot
```
Il attend ensuite que le nœud soit de nouveau accessible :
```yaml
- name: Wait for node to come back online
ansible.builtin.wait_for:
host: "{{ ansible_host }}"
port: "{{ opnsense_https_port }}"
state: started
timeout: 5400
delay: 30
when: needs_reboot
```
Enfin, il vérifie que la version du firmware ou la série du produit correspond à la cible attendue.
```yaml
- name: Check firmware version
ansible.builtin.uri:
url: "https://{{ opnsense_host }}/api/core/firmware/status"
method: GET
user: "{{ opnsense_api_key }}"
password: "{{ opnsense_api_secret }}"
force_basic_auth: true
validate_certs: false
register: _post_firmware_status
until: _post_firmware_status.json.product['product_' ~ firmware_target_kind] | default('unknown') == firmware_target_value
retries: 240
delay: 15
```
Cette vérification permet de confirmer de manière fiable que la mise à jour ou l'évolution de version a bien atteint la cible attendue.
## Mise à jour du nœud maître avec un snapshot Proxmox
Le nœud maître bénéficie de protections supplémentaires.
Comme il fonctionne sur Proxmox, le playbook crée un snapshot de la VM avant d'activer le mode maintenance CARP et de démarrer l'opération sur le firmware.
Pour cela, je crée un utilisateur et un token Proxmox dédiés à Semaphore :
```bash
pveum user add semaphore@pve
pveum user token add semaphore@pve opnsense -expire 0 -privsep 0
```
Je crée ensuite un rôle limité :
```bash
pveum role add SemaphoreOpnsenseUpdate -privs "\
VM.Audit \
VM.PowerMgmt \
VM.Snapshot \
VM.Snapshot.Rollback \
"
```
Le rôle est attribué uniquement à la VM OPNsense :
```bash
pveum aclmod /vms/122 -user semaphore@pve -role SemaphoreOpnsenseUpdate
```
J'aime cette approche, car Semaphore ne peut agir que sur la VM concernée par ce workflow. Il ne dispose pas de permissions étendues sur l'ensemble de l'environnement Proxmox.
Dans Semaphore, j'ajoute un autre groupe de variables pour les identifiants de l'API Proxmox :
- `PROXMOX_HOST`
- `PROXMOX_PORT`
- `PROXMOX_TOKEN_ID`
- `PROXMOX_USER`
- `PROXMOX_TOKEN_SECRET`
Pour utiliser les modules Proxmox, j'ajoute un fichier `requirements.yml` à côté du playbook :
```yaml
---
collections:
- name: community.proxmox
version: "2.0.0"
```
La collection Proxmox nécessite également la bibliothèque Python `proxmoxer`. J'ajoute donc un fichier `requirements.txt` à côté du fichier `docker-compose.yml` de Semaphore :
```text
proxmoxer>=2.3
```
Je le monte ensuite dans le conteneur Semaphore :
```yaml
volumes:
- /appli/docker/semaphore/requirements.txt:/etc/semaphore/requirements.txt
```
Après avoir redéployé Semaphore, le playbook peut créer le snapshot :
```yaml
- name: Take Proxmox VM snapshot
community.proxmox.proxmox_snap:
vmid: "{{ proxmox_vmid }}"
state: present
snapname: "{{ proxmox_snap_name }}"
description: "Pre-firmware-{{ firmware_action }}: {{ firmware_current_version }} → {{ firmware_target_value }}"
```
Si quelque chose échoue pendant la mise à jour du maître, le bloc de récupération restaure la VM depuis le snapshot créé avant la mise à jour et envoie une notification Ntfy de priorité élevée.
## Notification finale
Au début, j'utilisais trop d'assertions pour piloter la logique de notification. Cela fonctionne pour les échecs, mais ce n'est pas le bon modèle pour les situations normales comme l'absence de mise à jour disponible.
Le bloc de récupération doit uniquement gérer les véritables échecs. Les situations normales doivent atteindre la phase de notification finale.
La phase finale s'exécute sur `localhost` et compare les facts collectés sur les nœuds maître et de secours. Elle gère les deux cas suivants :
- Les deux nœuds ont effectué la même opération
- Chaque nœud possède un résultat différent
Le corps de la notification est généré à partir des variables d'hôte des nœuds maître et de secours :
```yaml
body: |
{% if same_operation %}
{% if m.skip_update | default(false) %}
Les deux nœuds sont déjà en {{ m.firmware_current_version }}, aucune action effectuée.
{% else %}
Cluster OPNsense : {{ m.firmware_current_version }} → {{ m.firmware_target_value }} ({{ m.firmware_status }})
{% endif %}
{% else %}
{% if m.skip_update | default(false) %}
MAÎTRE ({{ master }}) : déjà en {{ m.firmware_current_version }}, aucune action effectuée.
{% else %}
MAÎTRE ({{ master }}) : {{ m.firmware_current_version }} → {{ m.firmware_target_value }} ({{ m.firmware_status }})
{% endif %}
{% if b.skip_update | default(false) %}
SECOURS ({{ backup }}) : déjà en {{ b.firmware_current_version }}, aucune action effectuée.
{% else %}
SECOURS ({{ backup }}) : {{ b.firmware_current_version }} → {{ b.firmware_target_value }} ({{ b.firmware_status }})
{% endif %}
{% endif %}
```
La priorité et le tag de la notification changent également selon qu'une action a été effectuée ou que les deux nœuds sont déjà à jour.
J'obtiens ainsi un rapport utile sans transformer une exécution sans action en erreur.
## Workflow final
Le workflow terminé est divisé en quatre phases :
- Vérification du firmware sur tous les nœuds
- Mise à jour du nœud de secours sur TrueNAS
- Mise à jour du nœud maître sur Proxmox avec un snapshot
- Envoi d'une notification Ntfy
Le nœud de secours est mis à jour en premier. Le nœud maître est mis à jour en second, avec la création d'un snapshot Proxmox avant l'opération sur le firmware. L'état CARP est vérifié avant le début du workflow et le mode maintenance est utilisé pendant les mises à jour des nœuds.
Le playbook prend en charge les scénarios de mise à jour, d'évolution de version et de vérification au moyen du sondage Semaphore. Il sait également ignorer un nœud lorsqu'il n'y a rien à faire ou lorsque l'action demandée ne correspond pas à ce que signale OPNsense.
Plus important encore, le workflow s'exécute désormais de bout en bout et signale le résultat.
Le playbook Ansible est disponible [ici](https://github.com/Vezpi/Homelab/blob/main/ansible/opnsense/update_opnsense_ha_cluster.yml).
## Conclusion
Cette automatisation est partie d'une idée simple : ne plus mettre OPNsense à jour manuellement.
En pratique, le sujet s'est révélé plus intéressant qu'un simple appel à l'endpoint du firmware. Le playbook devait comprendre l'état HA, gérer différemment les mises à jour et les évolutions de version, mettre les nœuds à jour dans le bon ordre, protéger le maître hébergé sur Proxmox avec un snapshot et signaler l'état final sans considérer les exécutions sans action comme des échecs.
Le résultat s'intègre beaucoup mieux au reste de l'automatisation de mon homelab. Semaphore fournit un point d'entrée reproductible, Ansible gère la logique, OPNsense expose son état via son API, Proxmox fournit un point de restauration pour le maître et Ntfy m'indique ce qui s'est passé.
C'est une tâche de maintenance manuelle de moins à oublier, et un élément de plus du homelab capable de prendre soin de lui-même.
@@ -0,0 +1,453 @@
---
slug: automating-opnsense-update-ansible
title: Automating OPNsense HA updates with Ansible
description: Automating OPNsense HA updates in a homelab with Ansible, Semaphore UI, CARP checks, Proxmox snapshots and Ntfy notifications.
date: 2026-08-21
draft: false
tags:
- opnsense
- ansible
- semaphore-ui
- ntfy
- proxmox
categories:
- homelab
---
## Intro
In my homelab, most of the infrastructure is already redundant enough to tolerate maintenance, but the update process itself was still too manual. Proxmox was the first part I automated in that [post]({{< ref "post/20-automating-proxmox-update-ansible" >}}), and once that workflow was running from Semaphore on a schedule, the next logical target was OPNsense.
My OPNsense setup is an HA cluster with two nodes. The master node, `cerbere-head1`, runs on Proxmox. The backup node, `cerbere-head2`, runs on TrueNAS. That split is intentional, because I want the network to survive maintenance or outages on the Proxmox cluster.
The goal is simple: create an Ansible playbook able to update or upgrade the OPNsense HA cluster safely, in the right order, with checks before touching anything, and a notification at the end.
---
## Update strategy
OPNsense exposes an API that can be used to check system status, firmware status, services and CARP virtual IP state. Ansible drives the automation with API calls, while Semaphore UI is used as the controller to run the playbook manually or from a schedule. Ntfy is used for the final report and for failure notifications.
For the Proxmox hosted node, I also use the [community.proxmox Ansible collection](https://docs.ansible.com/projects/ansible/latest/collections/community/proxmox/index.html). This allows the playbook to create a VM snapshot before updating the master firewall node and roll back to it if needed.
The important detail is that both OPNsense nodes are not treated exactly the same. The backup node runs on TrueNAS, so the playbook updates it without taking a hypervisor snapshot. The master node runs on Proxmox, so the playbook takes a snapshot before starting the firmware operation.
---
## Creating an API user in OPNsense
To let Ansible interact with OPNsense, I create a dedicated user on the master node.
The user is called `automation`, with a scrambled password and only the privileges needed by the playbook:
- `Interface: Virtual IPS: Status`
- `System: Firmware`
- `System: Status`
- `Status: Services`
Because this is an HA cluster, the synchronization between both OPNsense nodes handles the user creation on the backup node.
After creating the user, I generate an API key.
![opnsense-user-create-api-key.png](images/opnsense-user-create-api-key.png)
The OPNsense API key is generated from the dedicated automation user.
The downloaded file contains the API `key` and `secret`. In the OPNsense UI, the key remains visible in the `ApiKeys` tab, but the secret does not.
I first test the API calls with Bruno from VS Code. Once the basic calls are working, I load the credentials into Semaphore.
---
## Preparing Semaphore UI
In Semaphore, I create a key store entry named `OPNsense automation`, containing the API key and secret.
Then I create an inventory for the OPNsense nodes:
```yaml
---
all:
children:
opnsense:
vars:
ansible_connection: local
children:
opnsense_backup:
hosts:
cerbere-head2:
ansible_host: 192.168.88.3
main_role: BACKUP
hypervisor: TrueNAS
opnsense_master:
hosts:
cerbere-head1:
ansible_host: 192.168.88.2
main_role: MASTER
hypervisor: Proxmox
proxmox_vmid: 122
```
The playbook runs locally from Semaphore and talks to each firewall through the OPNsense API.
I also create a variable group named `OPNsense automation API` with the API credentials and a few shared variables:
- `OPNSENSE_API_KEY`
- `OPNSENSE_API_SECRET`
- `opnsense_api_key`
- `opnsense_api_secret`
- `opnsense_https_port`
- `opnsense_host`
The HTTPS port is set to `4443`, and the host is built from the inventory address and this port.
Finally, I create the Semaphore task template.
![semaphore-new-template-task-opnsense-update.png](images/semaphore-new-template-task-opnsense-update.png)
The Semaphore task template used to run the OPNsense HA update playbook.
Before going further, I validate that Ansible could query both nodes:
```yaml
- name: Check node availability
ansible.builtin.uri:
url: "https://{{ opnsense_host }}/api/core/system/status"
method: GET
user: "{{ opnsense_api_key }}"
password: "{{ opnsense_api_secret }}"
force_basic_auth: true
validate_certs: false
```
At that point, the automation can reach both nodes and authenticate against the API.
---
## Making CARP maintenance usable from the API
One important part of the workflow is CARP maintenance mode.
Before updating a node, I want to put it into maintenance mode so the virtual IPs are not active on the node being updated. This is especially important when updating the master node, because the backup must take over cleanly before the update starts.
During API testing, the CARP maintenance endpoint returned `403 Forbidden`:
```text
POST https://{{opnsense_host}}/api/diagnostics/interface/carp_status/maintenance
```
The privileges were granted in the WebUI, so the issue looked related to the ACL definition. I manually adjusted the OPNsense ACL file by changing the API pattern to include the wildcard:
```xml
<pattern>api/diagnostics/interface/carp_status/*</pattern>
```
After rebooting the system, the endpoint returned a proper response.
I created a small [PR](https://github.com/opnsense/core/pull/10428) in the OPNsense project to fix this, and it was merged quickly into the `opnsense/core` `master` branch, my first contribution to that project.
![github-opnsense-pr-merged.png](images/github-opnsense-pr-merged.png)
The small OPNsense ACL fix was merged upstream.
Later, OPNsense [26.1.11](https://forum.opnsense.org/index.php?topic=52257.0) was released and included the fix. That allowed me to test the full playbook without relying on the manual ACL change.
---
## Designing the playbook workflow
The playbook follows a simple order:
- Check both nodes status
- Update the backup node first
- Update the master node second
- Send a final notification
That order is important. The backup node is updated first because the master is still active. Then, before updating the master, the playbook enables CARP maintenance mode to let the backup take over.
The playbook also supports different actions through a Semaphore survey.
![semaphore-opnsense-update-survey-action.png](images/semaphore-opnsense-update-survey-action.png)
The Semaphore survey lets me choose between check, update and upgrade.
This is needed because OPNsense does not expose updates and upgrades in exactly the same way. The variables for the target version and the reboot requirement differ between an update and an upgrade, so the playbook resolves those differences before deciding what to do.
---
## Firmware and CARP checks
The first phase runs on both OPNsense nodes.
It gathers information from several API calls:
- System status
- Firmware check
- Firmware status
- CARP virtual IP status
The playbook stores the result as facts that can be reused later:
```yaml
- name: Store node facts
ansible.builtin.set_fact:
firmware_action: "{{ opnsense_action | default('update')}}"
firmware_status: "{{ _firmware_status.json.status }}"
firmware_status_msg: "{{ _firmware_status.json.status_msg }}"
firmware_current_version: "{{ _firmware_status.json.product.product_version | default('unknown') }}"
firmware_product_series: "{{ _firmware_status.json.product.product_series | default('unknown') }}"
firmware_update_version: "{{ _firmware_status.json.upgrade_packages | selectattr('name', 'equalto', 'opnsense') | map(attribute='new_version') | first | default('') }}"
firmware_upgrade_version: "{{ _firmware_status.json.upgrade_major_version | default('unknown') }}"
firmware_upgrade_message: "{{ _firmware_status.json.upgrade_major_message | regex_replace('<[^>]+>', ' ') | regex_replace('\\s{2,}', ' ') | trim | default('') }}"
firmware_up_to_date: "{{ _firmware_status.json.status_msg == 'There are no updates available on the selected mirror.' }}"
needs_reboot: "{{ (_firmware_status.json.upgrade_needs_reboot == '1') if _firmware_status.json.status == 'upgrade' else (_firmware_status.json.needs_reboot == '1') }}"
total_vips: "{{ _vip_status.json.rowCount }}"
mismatched_vips: "{{ _vip_status.json.rows | rejectattr('status', 'equalto', main_role) | list | length }}"
in_maintenance: "{{ _vip_status.json.carp.maintenancemode }}"
```
Then it resolves whether the target is a regular version or a product series:
```yaml
- name: Resolve update-vs-upgrade specifics
ansible.builtin.set_fact:
firmware_target_kind: "{{ 'series' if firmware_status == 'upgrade' else 'version' }}"
firmware_target_value: "{{ firmware_upgrade_version if firmware_status == 'upgrade' else firmware_update_version }}"
```
This makes the rest of the playbook easier to manage. It can later check whether the node reached the expected `version` or `series` without duplicating the whole logic.
The first phase also validates a few conditions before proceeding:
- The node must not already be in CARP maintenance mode
- The VIPs must match the expected role
- At least one VIP must be managed
If one of these checks fails, the playbook aborts and sends a Ntfy notification.
## Handling skip logic properly
One of the trickiest parts is not the update itself, but deciding when not to update.
After the first successful test, the backup node is already updated while the master is not. The next run should not update the backup again if it is already at the version the master is targeting.
I add skip logic for that case:
```yaml
- name: Backup already updated
ansible.builtin.set_fact:
skip_update: true
firmware_status: "skipped"
delegate_to: "{{ groups['opnsense_backup'][0] }}"
delegate_facts: true
run_once: true
when: >-
(hostvars[groups['opnsense_backup'][0]].firmware_current_version
if hostvars[groups['opnsense_master'][0]].firmware_target_kind == 'version'
else hostvars[groups['opnsense_backup'][0]].firmware_product_series)
== hostvars[groups['opnsense_master'][0]].firmware_target_value
```
Then I generalize the skip behavior.
The playbook skips a node when:
- No updates are available
- An upgrade is available but the requested action is update
- An update is available but the requested action is upgrade
- The backup node is already at the target version or series of the master
This makes the final notification much cleaner, because a skipped node is not treated as an error. It is simply reported as no action needed.
## Updating the backup node
The backup node runs on TrueNAS, so this phase does not create a hypervisor snapshot.
The playbook enables CARP maintenance mode, triggers the firmware action, waits for the update to start, waits for the node to reboot if required, then waits for it to come back online.
The relevant part looks like this:
```yaml
- name: Trigger firmware {{ firmware_action }}
ansible.builtin.uri:
url: "https://{{ opnsense_host }}/api/core/firmware/{{ firmware_action }}"
method: POST
user: "{{ opnsense_api_key }}"
password: "{{ opnsense_api_secret }}"
force_basic_auth: true
validate_certs: false
```
If a reboot is required, the playbook waits for the HTTPS port to go down:
```yaml
- name: Wait for node to reboot after the {{ firmware_action }}
ansible.builtin.wait_for:
host: "{{ ansible_host }}"
port: "{{ opnsense_https_port }}"
state: stopped
timeout: 3600
when: needs_reboot
```
Then it waits for the node to come back:
```yaml
- name: Wait for node to come back online
ansible.builtin.wait_for:
host: "{{ ansible_host }}"
port: "{{ opnsense_https_port }}"
state: started
timeout: 5400
delay: 30
when: needs_reboot
```
Finally, it checks that the firmware version or product series matches the expected target.
```yaml
- name: Check firmware version
ansible.builtin.uri:
url: "https://{{ opnsense_host }}/api/core/firmware/status"
method: GET
user: "{{ opnsense_api_key }}"
password: "{{ opnsense_api_secret }}"
force_basic_auth: true
validate_certs: false
register: _post_firmware_status
until: _post_firmware_status.json.product['product_' ~ firmware_target_kind] | default('unknown') == firmware_target_value
retries: 240
delay: 15
```
That check is what gives the playbook a reliable confirmation that the update or upgrade actually reached the expected target.
## Updating the master node with a Proxmox snapshot
The master node is handled with more protection.
Because it runs on Proxmox, the playbook creates a VM snapshot before enabling CARP maintenance mode and starting the firmware action.
For this, I create a dedicated Proxmox user and token for Semaphore:
```bash
pveum user add semaphore@pve
pveum user token add semaphore@pve opnsense -expire 0 -privsep 0
```
Then I create a limited role:
```bash
pveum role add SemaphoreOpnsenseUpdate -privs "\
VM.Audit \
VM.PowerMgmt \
VM.Snapshot \
VM.Snapshot.Rollback \
"
```
The role is assigned only to the OPNsense VM:
```bash
pveum aclmod /vms/122 -user semaphore@pve -role SemaphoreOpnsenseUpdate
```
I like this approach because Semaphore can only operate on the one VM involved in this workflow. It does not get broad permissions on the whole Proxmox environment.
In Semaphore, I add another variable group for the Proxmox API credentials:
- `PROXMOX_HOST`
- `PROXMOX_PORT`
- `PROXMOX_TOKEN_ID`
- `PROXMOX_USER`
- `PROXMOX_TOKEN_SECRET`
To use the Proxmox modules, I added a `requirements.yml` next to the playbook:
```yaml
---
collections:
- name: community.proxmox
version: "2.0.0"
```
The Proxmox collection also requires the `proxmoxer` Python library, so I add a `requirements.txt` next to the Semaphore `docker-compose.yml`:
```text
proxmoxer>=2.3
```
Then I mounted it into the Semaphore container:
```yaml
volumes:
- /appli/docker/semaphore/requirements.txt:/etc/semaphore/requirements.txt
```
After redeploying Semaphore, the playbook can create the snapshot:
```yaml
- name: Take Proxmox VM snapshot
community.proxmox.proxmox_snap:
vmid: "{{ proxmox_vmid }}"
state: present
snapname: "{{ proxmox_snap_name }}"
description: "Pre-firmware-{{ firmware_action }}: {{ firmware_current_version }} → {{ firmware_target_value }}"
```
If something fails during the master update, the rescue block rolls the VM back to the pre-update snapshot and sends a high priority Ntfy notification.
## Final notification
At first, I used assertions too much to drive the reporting logic. That works for failures, but it is not the right model for normal cases like no updates available.
The rescue block should only handle real failures. Normal situations should reach the final notification phase.
The final phase runs on `localhost` and compares the facts collected from the master and backup nodes. It handles both cases:
- Both nodes had the same operation
- Each node had a different result
The notification body is generated from the master and backup host variables:
```yaml
body: |
{% if same_operation %}
{% if m.skip_update | default(false) %}
Both nodes already on {{ m.firmware_current_version }}, no action taken.
{% else %}
OPNsense cluster: {{ m.firmware_current_version }} → {{ m.firmware_target_value }} ({{ m.firmware_status }})
{% endif %}
{% else %}
{% if m.skip_update | default(false) %}
MASTER ({{ master }}): already on {{ m.firmware_current_version }}, no action taken.
{% else %}
MASTER ({{ master }}): {{ m.firmware_current_version }} → {{ m.firmware_target_value }} ({{ m.firmware_status }})
{% endif %}
{% if b.skip_update | default(false) %}
BACKUP ({{ backup }}): already on {{ b.firmware_current_version }}, no action taken.
{% else %}
BACKUP ({{ backup }}): {{ b.firmware_current_version }} → {{ b.firmware_target_value }} ({{ b.firmware_status }})
{% endif %}
{% endif %}
```
The notification priority and tag also change depending on whether an action is performed or both nodes are already up to date.
This gives me a useful report without turning a no-op run into an error.
## The final workflow
The finished workflow is split into four phases:
- Firmware check on all nodes
- Update the backup node on TrueNAS
- Update the master node on Proxmox with a snapshot
- Send a Ntfy notification
The backup node is updated first. The master node is updated second, with a Proxmox snapshot taken before the firmware action. CARP status is checked before the workflow starts, and maintenance mode is used during node updates.
The playbook can handle update, upgrade and check scenarios through the Semaphore survey. It also knows when to skip a node because there is nothing to do or because the requested action does not match what OPNsense reports.
Most importantly, the workflow now completes end to end and reports the result.
The Ansible playbook can be found [here](https://github.com/Vezpi/Homelab/blob/main/ansible/opnsense/update_opnsense_ha_cluster.yml).
## Conclusion
This automation started as a simple idea: stop updating OPNsense manually.
In practice, it became more interesting than just calling the firmware endpoint. The playbook needed to understand the HA state, handle updates and upgrades differently, update nodes in the right order, protect the Proxmox hosted master with a snapshot, and report the final state without treating normal no-op cases as failures.
The result is a workflow that fits much better with the rest of my homelab automation. Semaphore gives me a repeatable entry point, Ansible handles the logic, OPNsense exposes the state through its API, Proxmox provides a rollback point for the master node, and Ntfy tells me what happened.
It is one less manual maintenance task to forget, and one more piece of the homelab that can take care of itself.