Monitoring Systems Engineer – Middle

Remote middle English B1 2 months ago full-time
Apply now →

✓ Drop your CV once, then continue to the employer's application form. Your profile stays here for every recruiter hiring on igamingjobs.

Role in brief

SOFTSWISS is looking for a Middle Monitoring Systems Engineer to improve their monitoring ecosystem for critical services. Applicants should have experience in systems engineering and a strong background in Linux and containerization technologies.

LinuxDockerKubernetesBashPythonGoPostgreSQLAI tools

About the role

In this role, you will be responsible for designing, maintaining, and enhancing the monitoring and observability systems at SOFTSWISS. Your work will ensure the reliability and performance of critical services, and you will collaborate with various teams to meet their monitoring needs.

You will also handle incident troubleshooting, maintain detailed documentation, and contribute to best practices in AI usage within monitoring workflows. Success in this position means effectively implementing solutions that enhance system visibility and reduce noise in alerts.

Skills that matter here

  • Linux: A strong understanding of Linux-like operating systems is essential for maintaining the monitoring infrastructure.
  • Docker: Experience with containerization technologies like Docker is required for setting up and managing monitoring systems.
  • Kubernetes: Knowledge of orchestration tools like Kubernetes is necessary for scaling monitoring solutions.
  • Bash: Scripting in Bash will be used to automate monitoring tasks and enhance operational efficiency.
  • Python: Development experience in Python is important for creating custom monitoring solutions.
  • PostgreSQL: A basic understanding of PostgreSQL is preferred for database-related monitoring tasks.

Who this role suits

  • Detail-oriented individuals who can manage multiple tasks effectively.
  • Team players who thrive in collaborative environments and can communicate technical needs.
  • Candidates with a proactive approach to problem-solving and a strong interest in automation.
  • Fluent in both English and Russian, with a critical eye for evaluating AI tools.

From the employer

  • Offering on-duty service coverage, encompassing day and night shifts.
  • Addressing incidents by troubleshooting and resolving issues, even seeking assistance from third-party or vendor support when necessary.
  • Directing issues or queries to the relevant department as needed.
  • Keeping detailed records and documentation of current infrastructure challenges and Root Cause Analyses (RCAs).
  • Contribute to safe and effective internal practices for AI usage in monitoring and incident response workflows.
  • Collaborating with other teams to understand and define their monitoring needs, then implementing the right solutions.
  • Setting up and adjusting the monitoring/observability systems for various teams.
  • Designing and tweaking alerts and dashboards to suit specific needs.
  • Refining alerts to reduce irrelevant notifications and increase their significance.
  • Enhancing dashboards for better clarity, understanding, and a more comprehensive view.
  • Building and sustaining connections between the monitoring systems and other platforms like Jira, Opsgenie, etc. when required.
  • Establishing and updating a Knowledge Base, covering system configurations, alert processes, troubleshooting guidelines, and user manuals.
  • Staying updated with the newest trends and best practices to continuously uplift our organization's monitoring capabilities.
  • Identify opportunities to automate repetitive monitoring and support tasks, including with AI-assisted approaches where suitable.
  • Minimum of 3 years experience as a Systems Engineer, SRE, DevOps, or Monitoring Support Engineer (L2+).
  • Good understanding of Linux-like operating systems (Debian-based).
  • Experience with containerization, virtualization, and orchestration (LXC/LXD, Docker, Kubernetes).
  • Development experience in any scripting language (Bash, Python, Go, etc) and familiarity with REST API.
  • Knowledge of basic database concepts (experience with PostgreSQL is preferable), including transactions and WAL.
  • English proficiency at an Intermediate (B1) level or higher.
  • Russian proficiency at an Upper-Intermediate (B2) level or higher.
  • Practical interest in using AI-assisted tools for troubleshooting, automation, documentation, and operational efficiency.
  • Ability to critically evaluate AI-generated output and validate it before using it in production environments.
  • Understanding of the risks and limitations of AI usage in infrastructure and production operations.
  • Private health insurance
  • Sports benefits
  • Comprehensive Mental Health Program
  • Free English lessons (online)
  • Local language courses
  • Paid time off
  • Maternity leave support
  • Referral program rewards
  • Upskilling, internal workshops, and participation in professional conferences and corporate events

Questions about this role

Is this position remote?

Yes, this role is remote.

What is the seniority level of this position?

This is a middle-level position.

How can I apply for this role?

Candidates can apply through the SOFTSWISS website.

Apply now →

✓ Drop your CV once, then continue to the employer's application form. Your profile stays here for every recruiter hiring on igamingjobs.

Similar jobs