9+ High Performance Computing (HPC) Jobs in India
Apply to 9+ High Performance Computing (HPC) Jobs on CutShort.io. Find your next job, effortlessly. Browse High Performance Computing (HPC) Jobs and apply today!

Job Overview
We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms. This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.
This is not a DevOps-focused role. We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms.
The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.
Key Responsibilities & Required Skills
Linux & Bare Metal Infrastructure
- Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)
- Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
- Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
- Strong understanding of server hardware, including:
- BIOS/UEFI
- RAID controllers
- Firmware management
- iLO/iDRAC/IPMI
- NICs and SmartNICs
- HBA cards
- Hardware diagnostics and troubleshooting
- Experience designing, implementing, and supporting enterprise Linux infrastructure at scale
AI Factory & GPU Infrastructure
- Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads
- Understanding of NVIDIA GPU technologies including:
- A100, H100, H200, B200, or equivalent GPU platforms
- NVIDIA DGX and OEM GPU servers
- GPU provisioning and lifecycle management
- GPU monitoring and performance optimization
- Knowledge of AI Factory architecture and infrastructure requirements
- Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads
- Understanding of:
- GPU resource allocation and scheduling
- Multi-GPU systems
- GPU networking requirements
- High-bandwidth, low-latency infrastructure design
- Familiarity with NVIDIA ecosystem technologies such as:
- CUDA
- NCCL
- GPUDirect Storage
- NVIDIA Fabric Manager
- NVIDIA Base Command (preferred)
Enterprise Storage & Data Platforms
- Advanced Linux storage administration:
- LVM
- XFS, EXT4
- NFS
- iSCSI
- Fibre Channel SAN
- Multipath I/O
- Strong hands-on experience with Ceph, including:
- Cluster architecture
- MON, OSD, MDS
- RBD, CephFS, RGW
- Capacity planning
- Performance tuning
- Failure recovery
- Experience with high-performance AI storage platforms such as:
- WEKA
- VAST Data
- Dell PowerScale
- Pure Storage FlashBlade
- NetApp
- Understanding of:
- NVMe-over-Fabrics (NVMe-oF)
- RDMA
- GPUDirect Storage
- Parallel file systems
- AI data pipelines
Networking & Infrastructure
- Strong networking knowledge:
- Bonding
- VLANs
- Routing
- MTU optimization
- DNS
- DHCP
- Experience with high-performance data center networking:
- 100G/200G/400G Ethernet
- RoCE
- RDMA
- Spine-Leaf architectures
- Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
- Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting
Operations & Reliability
- Experience with high availability, clustering, and disaster recovery
- Strong troubleshooting skills across:
- Linux operating systems
- Hardware platforms
- GPU infrastructure
- Networking
- Enterprise storage
- Experience supporting mission-critical production environments
- Bash and Python scripting for automation and operational efficiency
- Experience creating operational documentation, runbooks, and infrastructure standards
Nice to Have
- Kubernetes infrastructure (especially AI/ML and GPU integration)
- KVM, VMware, OpenShift Virtualization, or similar virtualization platforms
- Ansible automation
- NVIDIA Base Command Manager
- Slurm or HPC workload schedulers
- Observability and monitoring platforms (Prometheus, Grafana, OpenTelemetry)
- Data Center Infrastructure Management (DCIM) tools
- IPAM solutions
- AWS, Azure, or hybrid cloud exposure
Greeting from Nam info!
We have a job opportunity for HPC Expert ClearML & Confidential Computing , interested candidate & who is available for Final round F2F interview in Bangalore,
Role : HPC Expert ClearML & Confidential Computing
Exp : 5+years
Work Location : Bangalore
Notice period: Immediate - 15Days
Key Responsibilities
· Administer ClearML Server, ClearML Agents, execution queues, projects, users, roles, experiment tracking, pipelines, datasets, artifacts, and model registry.
· Configure ClearML Agents on CPU and GPU worker nodes and integrate them with Slurm, PBS Professional, Kubernetes, or equivalent HPC execution platforms.
· Support GPU-based AI/ML workloads using NVIDIA drivers, CUDA, NCCL, UCX, and containerized environments.
· Maintain secure container execution using Apptainer/Singularity, Docker, Enroot, Pyxis, or Kubernetes.
· Implement confidential-computing controls using technologies such as AMD SEV-SNP, Intel TDX, NVIDIA Confidential Computing, Secure Boot, TPM, and remote attestation where supported.
· Integrate authentication, RBAC, TLS certificates, secrets management, and audit controls for ClearML and confidential workloads.
· Monitor ClearML services, agents, queues, GPU utilization, task failures, scheduler integration, and platform health.
· Automate deployment, configuration, monitoring, and troubleshooting using Python, Bash, Ansible, and Git.
Software Engineer, Low-Latency Systems
- Employment Type: Full-time
- Experience Level: senior-level (7–10 years)
About the Role
We are hiring a Software Engineer, Low-Latency Systems to design and optimize the core infrastructure powering our algorithmic trading systems. In this role, you will work on latency-critical execution paths where nanoseconds, cache lines, memory layout, and network behavior matter.
This is a hands-on engineering position for someone who enjoys building high-performance systems and reasoning deeply about correctness, throughput, and tail latency. Prior trading domain experience is helpful but not required—we value engineering depth and systems thinking above all else.
What You’ll Do (Responsibilities)
- Build Core Infrastructure: Design, develop, and maintain low-latency components including order routing, market data handling, and execution pipelines.
- Optimize Performance: Profile and optimize critical code paths to minimize throughput and tail latency.
- Collaborate Across Teams: Work closely with quant and trading teams to translate complex strategy requirements into highly efficient infrastructure primitives.
- Drive System Design: Contribute to architectural decisions around threading models, memory layout, and network stack configurations.
- Ensure Reliability: Improve observability and operational performance across trading infrastructure. Participate in on-call rotations, incident response, and post-mortems to keep systems running smoothly.
What We’re Looking For (Requirements)
- Experience: 7 to 10 years of professional experience in systems engineering, with a demonstrable focus on low-latency systems or high-performance computing (HPC).
- Language Proficiency: Strong, production-level proficiency in Rust and/or C++.
- Systems Depth: Comfort reasoning about memory management, lock-free data structures, compiler behavior, and CPU-level performance.
- Tooling: Experience using Linux performance tooling such as perf, flamegraphs, strace, or similar tools.
- Networking Fundamentals: Solid understanding of network stack behavior, including TCP, UDP, multicast, and kernel bypass.
- Problem Solving: Ability to debug complex production issues and optimize systems under real-world constraints.
Nice to Have (Bonus Points)
- Prior exposure to trading systems, market data feeds, or exchange connectivity.
- Familiarity with financial market protocols (e.g., FIX, ITCH, OUCH).
- Experience with low-latency networking technologies like DPDK, RDMA, or kernel bypass.
- Familiarity with co-location environments and latency-sensitive infrastructure.
Culture & Fit
We are looking for an engineer who takes ownership, thrives in ambiguous and fast-moving environments, and holds an incredibly high bar for correctness and performance. If you love drilling down into the lowest levels of software to squeeze out maximum efficiency, we want to hear from you.
As a Lead Solutions Architect at Aganitha, you will:
* Engage and co-innovate with customers in BioPharma R&D
* Design and oversee implementation of solutions for BioPharma R&D * Manage Engineering teams using Agile methodologies
* Enhance reuse with platforms, frameworks and libraries
Applying candidates must have demonstrated expertise in the following areas:
1. App dev with modern tech stacks of Python, ReactJS, and fit for purpose database technologies
2. Big data engineering with distributed computing frameworks
3. Data modeling in scientific domains, preferably in one or more of: Genomics, Proteomics, Antibody engineering, Biological/Chemical synthesis and formulation, Clinical trials management
4. Cloud and DevOps automation
5. Machine learning and AI (Deep learning)
● Design and deliver scalable web services, APIs, and backend data modules.
Understand requirements and develop reusable code using design patterns &
component architecture and write unit test cases.
● Collaborate with product management and engineering teams to elicit &
understand their requirements & challenges and develop potential solutions
● Stay current with the latest tools, technology ideas, and methodologies; share
knowledge by clearly articulating results and ideas to key decision-makers.
Requirements
● 3-6 years of strong experience in developing highly scalable backend and
middle tier. BS/MS in Computer Science or equivalent from premier institutes
Strong in problem-solving, data structures, and algorithm design. Strong
experience in system architecture, Web services development, highly scalable
distributed applications.
● Good in large data systems such as Hadoop, Map Reduce, NoSQL Cassandra, etc.. Fluency in Java, Spring, Hibernate, J2EE, REST Services Ability to deliver code
quickly from given scenarios in a fast-paced start-up environment.
● Attention to detail. Strong communication and collaboration skills.
