
Real job — pulled straight from AIPHORIA’s careers page · Verified September 22, 2026 · No reposts.
Job description
AIPHORIA is hiring a Platform Operations Engineer — a full-time, based in Argentina role. Apply directly on AIPHORIA's careers page below.
Platform Operations Engineer
Location: Argentina (AR)
Experience Level: Senior
Description
We're growing Platform Operations and building a system that can detect failures automatically, run proven recovery procedures, and bring engineers in when their expertise is needed. Our goal is a platform that runs reliably with minimal manual work.
To get there, we're looking for a Platform Operations Engineer who will learn how the platform works, investigate failures, restore services, and help turn practical experience into runbooks, checks, and automation.
If you feel you've reached a ceiling in technical support, NOC, presales, systems administration, or QA and want to expand your engineering skills, this role could be your next step. You'll build on your existing experience, work with production systems, and solve problems alongside DevOps engineers, SREs, and developers.
In this role, you'll shape how daily work with the platform gets done. Every incident you work through becomes the basis for an improvement: sharper monitoring, a clearer recovery procedure, or a script that handles the job automatically next time.
Requirements
- Linux and the command line. Navigate the filesystem, read and filter files, inspect processes, available resources, and environment variables.
- Application troubleshooting. Reproduce a problem, gather facts, test a hypothesis, and narrow the search. Know how to tell an application bug from a misconfiguration, an unavailable dependency, or a lack of resources.
- Working with logs. Find events by timestamp, request ID, or other markers, correlate entries across several services, and pull out the data that's actually useful for diagnosis.
- Monitoring and metrics. Hands-on experience with Grafana or similar tools. Understand availability, error rate, response time, and resource consumption, notice changes, and connect them to incidents.
- Networking and HTTP/APIs. Understand the basics of DNS, IP addressing, ports, and HTTP. Check whether a service is reachable, send a request with curl or Postman, and work out a status code, timeout, or authorization error.
- Application configuration. Read YAML and JSON, check environment variables, and compare settings across environments. Understand how configuration affects connections and application behavior.
- Executing technical procedures. Follow a runbook, verify preconditions before running a command and the result afterward. Know the limits of your access, when to stop, and how rollback and escalation work.
Nice to have
- Experience with Docker and Kubernetes: inspecting container and pod state, events, and logs via kubectl.
- Familiarity with Helm, ConfigMaps, Secrets, and how configuration is organized in Kubernetes.
- Basic Git skills and an understanding of how applications are built, deployed, and rolled back through CI/CD.
- Ability to read, modify, and write small Python or Bash scripts.
- Experience automating processes in n8n: API and webhook integrations, event handling, and triggering automated actions.
- Experience using Claude Code, Codex, or similar AI tools to explore code, analyze errors, write scripts, and prepare documentation, along with the judgment to verify what they suggest and what their output actually does.
- Basic SQL for checking data and tracing the source of errors.
- Experience with microservices, message queues, telephony, or AI inference services.
Soft skills
- Technical curiosity. A drive to understand how the system is built and why it behaves the way it does, to pick up new tools, and to put what you learn to use.
- English. Read technical documentation, handle written communication, and discuss technical matters with colleagues and users.
- Asking good questions. Clarify symptoms, context, and the expected outcome to get to the heart of a problem faster.
- Clear communication. Describe a problem in a structured way, pass on context when escalating, and explain the current status to people with different levels of technical background.
- Independence and ownership. Organize your own work, track next steps, and follow tasks through to a confirmed result.
- Knowing when to pull people in. Recognize the limits of your knowledge and access, and bring in the right specialists in good time.
- Prioritization and composure during incidents. Assess user impact, choose the order of actions, and stay consistent when several things are happening at once.
- Initiative and teamwork. Suggest improvements, share what your investigations turn up, take feedback, and help turn recurring problems into proven procedures and automation.
Responsibilities
- Detect and investigate problems. Respond to alerts, E2E check results, and requests coming in via Telegram, Slack, and email. Clarify symptoms, assess user impact, analyze logs, service state, and configuration, test hypotheses, and pin down the failure.
- Restore the platform. Execute proven runbooks within the access you've been granted, including restarts, redeployments, and failover to backup resources. Confirm that the product is working again.
- Escalate to specialist teams when needed. Bring in DevOps, SRE, SIP engineers, and developers promptly when their expertise is required, when no suitable runbook exists, or when a procedure didn't produce the expected result. Hand over the facts you've gathered, the actions you've taken, and your diagnostic findings.
- Own the incident from the first signal to closure. Record the timeline and the outcome of each action, keep the status current, track next steps, and keep everyone involved informed of progress.
- Write postmortems and turn findings into improvements. Work through causes and consequences with the people involved, and assess how well diagnosis and recovery went. Turn the results into bug reports and fix tasks, write and update runbooks, and help automate checks and recurring operations.
What we offer
- The team has built award-winning AI products for tech corporations - devices, voice assistants, products that are actually in the world
- Cutting-edge tech stack: Speech Technologies, NLP, Generative AI (LLMs, diffusion models), voice-first agentic architecture with privacy-first and on-premises deployment
- High engineering bar and real ownership - the team cares about what actually works in production, not what looks good in a demo, and you'll see the impact of your work directly
- Fast career progression - a senior-heavy team and a high volume of real problems means you grow faster than you would anywhere else
- Startup pace with enterprise stability - real clients, real revenue, no bureaucracy
- Fully remote across Europe
- 21 vacation days + public holidays + 5 sick days
- Private English lessons via Preply
Get Platform Operations Engineer jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs




Frequently asked questions
What skills are required for Platform Operations Engineer at AIPHORIA?
The required skills for Platform Operations Engineer at AIPHORIA include: Linux, Bash, Python, Grafana, DNS, HTTP, API, JSON, Git, Docker, Kubernetes, Helm, CI/CD, SQL, Microservices, AI, LLM, NLP, Postman.
What is the seniority level for Platform Operations Engineer at AIPHORIA?
Platform Operations Engineer at AIPHORIA is a Senior level position.
How do I apply for Platform Operations Engineer at AIPHORIA?
You can view the full description and apply for Platform Operations Engineer at AIPHORIA on EchoJobs: https://echojobs.io/job/aiphoria-platform-operations-engineer-j1x3m.