AI Engineer
Unicloud
Batch: 2023/2024/2025/2026. Responsibilities:
- Deploy, update, restart, and manage GPU-based AI applications and services.
- Monitor AI services, APIs, application logs, performance, and availability.
- Monitor GPU health, utilization, VRAM, temperature, power usage, and processes.
- Monitor CPU, RAM, disk, and other infrastructure resources.
- Troubleshoot application crashes, deployment failures, GPU issues, and service-related problems.
- Perform Linux server health checks and troubleshooting.
- Manage and troubleshoot Docker containers for AI applications.
- Maintain deployment configurations and operational procedures.
- Develop basic Python/Shell scripts for monitoring, health checks, and automation.
Requirements:
- Strong troubleshooting and analytical skills with a structured approach to problem-solving.
- Comfortable working with Linux, logs, monitoring tools, and command-line environments.
- Hands-on exposure to NVIDIA GPUs and GPU monitoring tools.
- Familiarity with AI/LLM applications and AI serving platforms such as vLLM.
- Working knowledge of Docker and Kubernetes.
- Exposure to Prometheus and Grafana.
- Basic understanding of Git/version control workflows.
- Strong interest in AI, GPU infrastructure, and LLM technologies.
- Ability to take ownership of assigned tasks and work independently.
- Willingness to continuously learn and grow in AI infrastructure and operations.
Note: Share your CV.
Apply via EmailOpen in CarrerliftTry more jobs free
Posted 2026-08-25