Job Description
Description: Perform general server and cluster administration tasks; monitoring and optimizing system performance and reliability, operational workflow development, and managing enhancements/upgrades and providing various levels of support; standing up and maintaining Kubernetes cluster, provisioning, and labeling new systems for efficient scheduling, orchestrating containers and storage for optimal compute and memory, storage, and network I/O. The CW will be working from an SLA-driven ticket queue with a DevOps support team that administers Linux and Windows environments for AI software development. Fast task turnaround is the priority.
Key Responsibilities: • Develop and maintain system documentation for Lab/Data Center configuration and customizations.
• Designs, develops and maintains infrastructure and templates for efficient software deployment and maintenance.
• Analyze performance on systems, generate reports, propose solutions to problems; implements performance improvements in collaboration with Architects, Engineers, and Network Engineer to ensure system efficiency.
• Provides hardware, OS, kernel, and application level support for development system environments.
• Perform software, kernel upgrade, patch updates, and data migration, for server software.
• Apply configuration and tuning standards in accordance with Linux and Windows recommendations and Infotree's requirements.
• Develop and maintain system documentation for server/software configuration and customizations.
• Manage/automate K8s interactive and queued pods & containers
• Design containers for optimal use of underlying hardware
• Replicate/mirror networked storage to local caches
• Provide Tier 2 and Tier 3 level technical support.
• Hardware/Server landing and OS deployment to documented specifications (rack, wire, configure) within Hillsboro labs as needed
• Linux and Windows Administration and troubleshooting (Multiple Distro)
• Asset management and maintenance using standard tools
• Use ticketing & instant messaging systems to communicate with customers and track work
• Communicate with email, IM, JIRA, and phone with our internal customer groups for tickets
• Comfortable reviewing and modifying BIOS, RAID, and BMC/IPMI settings
Required Skills and Experience: • Minimum 5+ years IT related experience
• Minimum Education: Bachelor's Degree in CS or related field for US candidates
• Linux OS administration (CentOS/Rocky/Red Hat & Ubuntu/Debian) administration including installation, configuration, monitoring, upgrade and Pkg installation, and user and group management.
• Knowledge of computer diagnostics and installation, including hardware, driver, software troubleshooting, and networking
• Server HW administration experience: Familiarity with network interface card installation, switches, wiring, rack management
• Linux Shell scripting experience (to perform commands/scans across clusters - distributed servers)
• Experience standing up, administrating K8s clusters
• Experience with Jenkins or similar CI/CD tool
• Experience with Docker container creation and deployment
• Communication skills (working with team members and internal customers)
• Detail oriented and able to follow process/instructions
• Ability to work in a fast-paced environment and offer effective solutions under tight deadlines
• Strong problem-solving and root cause analysis skills
• Follow written and/or Verbal instructions for custom operating system and application installs Maintaining and auditing Lab assets and routine inventory control
• Capable of learning and using the custom-built asset and inventory tracking software (training will be provided)
Helpful Skills: • Experience with datacenter orchestration tools such as ansible, MAAS, OpenStack
• Linux network stack experience
• Ability to lift 1U to 4U rack-mounted servers up to 40 pounds and familiarity with the use of Rack Jacks and general data center safety procedures- Experience with git
Job Tags
Local area, Remote work