7 hours, 8 minutes ago

Monitoring & Observability Operator

1. Presentation of ETNIC

ETNIC (Entreprise pour les Technologies de l’Information et de la Communication) is the IT operator of the Fédération Wallonie-Bruxelles. As a public interest organization, ETNIC's mission is to design, develop, maintain, and evolve information systems and technological infrastructures serving the administrations and institutions of the FWB.

As a central player in the digital transformation of the Belgian French-speaking public sector, ETNIC operates in various fields such as:

  • Management of IT infrastructures (networks, security, data centers, cloud),
  • Development of custom business applications,
  • Support for digital projects (functional analysis, UX/UI, project management),
  • Cybersecurity and data protection,
  • User support and training.

With a constant focus on innovation, performance, and public service, ETNIC regularly collaborates with external partners to strengthen its teams through IT consultancy missions. These collaborations take place within an ethical, professional framework, oriented towards quality and the concrete impact of the delivered solutions.

2. Mission

Zabbix – Administration and Operation

  • Master the entire Zabbix environment: creation and maintenance of dashboards, network maps, triggers, items, and templates
  • Configure and adjust alert thresholds (triggers) according to business and technical needs, including via user macros and dependent triggers to reduce noise
  • Create and maintain reusable monitoring templates (items, macros, LLD, low-level discovery)
  • Design and maintain items of type Zabbix Agent, SNMP, JMX, HTTP Agent, and External Checks as per use cases
  • Develop custom discovery scripts (LLD) for the auto-discovery of equipment and services
  • Diagnose false positives, noisy alerts, and collection anomalies (proxy, pollers, timeouts) in coordination with technical teams under the supervision of the manager
  • Ensure daily monitoring of active alerts and their qualification (real incident vs noise)
  • Configure and maintain Zabbix ↔ ServiceNow integrations (automatic incident escalation) and Zabbix ↔ notification tools integrations
  • Administer the distributed Zabbix architecture (proxies, servers, high availability) and ensure its performance
  • Participate in the evolution of the monitoring architecture (proxies, SNMP v2/v3 integrations, VMware API, NetFlow, etc.)
  • Perform root cause analysis (RCA) on complex multi-layer incidents (network, system, application, storage) as requested by technical teams
  • Master the Zabbix Services module (Business Service Monitoring): build hierarchical service trees, aggregate statuses (SLA/SLO) from triggers, calculate and monitor service availability rates by business service
  • Model end-to-end application services by associating technical components (hosts, triggers) with relevant business services, for a user/business-oriented view rather than purely infrastructure
  • Configure problem tags and state propagation rules for accurate mapping between technical incidents and service impacts
  • Use SLA reports generated by the Services module to feed monthly reporting and availability reviews

Oh Dear – Website Monitoring

  • Set up and maintain technical monitoring of websites on Oh Dear (availability, SSL certificates, performance, uptime, broken links, mixed content)
  • Configure alerts (email, webhooks, third-party integrations) and check daily the status of monitored sites
  • Use the Oh Dear REST API for data extraction and integration with reporting tools (Power BI, a plus)
  • Ensure follow-up on detected incidents and escalate them to the relevant teams, including impact analysis

ServiceNow – Ticket and Incident Management

  • Daily check incidents created in ServiceNow related to monitoring and ensure their technical qualification
  • Handle all requests from IT teams arriving via our request catalog in ServiceNow and by email
  • Qualify, prioritize (according to impact/urgency), and process monitoring-related requests received via ServiceNow or by email
  • Contribute to the reliability of the CMDB in connection with discovery processes (VMware, network) and resolution of identification conflicts
  • Ensure rigorous follow-up until ticket closure, documenting actions taken

Observability

  • Participate in the continuous improvement of IT service observability by combining information from monitoring, logs, events, and supervision tools
  • Contribute to the definition and evolution of health, availability, and performance indicators for business services
  • Correlate alerts, metrics, events, and technical data to facilitate rapid identification of incident causes
  • Participate in reducing operational noise by improving monitoring rules, alert thresholds, and correlation mechanisms
  • Collaborate with infrastructure, network, applications, and integration teams to improve end-to-end visibility of critical services

Reporting

  • Identify recurring incidents and trends observed through monitoring tools and operational reports
  • Produce a monthly monitoring report (availability, incidents, trends, corrective actions, KPI/SLA) for beneficiaries and management and upon request
  • Automate report production via scripts and APIs (Zabbix API, Oh Dear API, ServiceNow API) to ensure reliability and speed up reporting
  • Build and maintain management dashboards (Power BI, a plus, or equivalent) for management and technical teams if needed

Network, API & Automation

  • Diagnose connectivity and collection incidents (SNMP, syslog, NetFlow, routing, ACL, VLAN, firewall)
  • Develop and maintain automation scripts (Python/Bash) for monitoring checks, LLD discovery scripts, and data processing (e.g. export/processing CSV/Excel)
  • Consume and integrate REST APIs between the different ecosystem tools (Zabbix, ServiceNow, Oh Dear, Power BI if needed)
  • Contribute to application integration between systems (data flows, JSON/XML formats) to ensure end-to-end reliability of the monitoring chain
  • Document technical procedures and diagnostics (diagrams, operational guides)

DevOps Culture

  • Apply CI/CD best practices for versioning and deployment of monitoring configurations (templates, dashboards)
  • Use Git for version control of scripts and configurations
  • Contribute to the automation of the monitoring infrastructure through Infrastructure as Code approaches (Ansible, Terraform)
  • Participate in the containerization of monitoring tools (Docker) if necessary
  • Maintain technical and operational documentation (wiki) as well as knowledge base (KB) articles
  • Ensure knowledge transfer to support and operations teams

As a consultant, you are subject to the same working conditions as our internal staff, meaning a hybrid mode combining onsite and remote work, with a mandatory minimum of 50% presence in our offices.

The mission is exclusively in French.

3. Activities

Ensure the administration, operation, and evolution of monitoring and observability tools, under the supervision of the Monitoring & CMDB manager. Activities include:

  • Administer Zabbix: creation and maintenance of dashboards, network maps, items, triggers, and templates, automation of discovery, and optimization of the monitoring architecture
  • Model and monitor end-to-end business and application services, notably via the Zabbix Services module, to assess incident impact and monitor availability and SLA/SLO commitments
  • Ensure website monitoring with Oh Dear: availability, SSL certificates, performance, and anomalies
  • Qualify and process alerts, incidents, and monitoring requests in ServiceNow, ensure their follow-up until closure, and contribute to CMDB reliability
  • Improve observability by correlating metrics, logs, and events, reducing unnecessary alerts, and contributing to incident diagnosis with technical teams
  • Develop integrations between tools and automate controls, data collection, and reporting using REST APIs and Python/Bash scripts
  • Produce availability and performance reports, analyze recurring incidents, and maintain management dashboards
  • Apply Git, CI/CD, and Infrastructure as Code practices, maintain technical documentation, and ensure knowledge transfer to support and operations teams

Apply for this Job

This position was originally posted on Pro Unity.

It is publicly accessible, and we recommend applying directly through the Pro Unity website instead of going through third party recruiters.

Newsletter signup illustration