Comments (11)
Yes. It is based on HPA. You can use scale metric as well for HPA to pick up automatically.
from training-operator.
@erictanjn The problem in GPU is that the resource metrics are CPU related which do not capture the GPU utilization, how can we adjust replicas then ?
Some custom GPU metrics can be used for HPA, such as DCGM
from training-operator.
hi,
I noticed HPA use crd's spec.replicas as currentReplicas but there is not .Spec.Replicas in pytorchjob crd. So I'm a little bit confused how it works. How HPA controller get the currentReplicas of a pytorchjob?
the code in HPA controller as following:
currentReplicas := scale.Spec.Replicas
from training-operator.
The HPA controller will obtain the PyTorchJob replicas from .spec.pytorchReplicaSpecs.Worker.replicas
:
from training-operator.
Hi @erictanjn , can you share some detail about your user case in detail plz, AFAK, HPA is more like CPU oriented, do you work in GPU case and how you expect it works in production ?
Thanks in advance.
from training-operator.
The HPA controller will obtain the PyTorchJob replicas from
.spec.pytorchReplicaSpecs.Worker.replicas
:
do I need to do some work in HPA controller? Or it will obtain the PyTorchJob replicas from .spec.pytorchReplicaSpecs.Worker.replicas
automatically?
from training-operator.
Hi @erictanjn , can you share some detail about your user case in detail plz, AFAK, HPA is more like CPU oriented, do you work in GPU case and how you expect it works in production ?
Thanks in advance.
I use it in GPU case. So I am a little bit worried it cannot get the PyTorchJob replicas automatically.
from training-operator.
@erictanjn The problem in GPU is that the resource metrics are CPU related which do not capture the GPU utilization, how can we adjust replicas then ?
from training-operator.
The HPA controller will obtain the PyTorchJob replicas from
.spec.pytorchReplicaSpecs.Worker.replicas
:
do I need to do some work in HPA controller? Or it will obtain the PyTorchJob replicas from
.spec.pytorchReplicaSpecs.Worker.replicas
automatically?
It's automatically.
from training-operator.
This issue has been automatically marked as stale because it has not had recent activity. It will be closed if no further activity occurs. Thank you for your contributions.
from training-operator.
This issue has been automatically closed because it has not had recent activity. Please comment "/reopen" to reopen it.
from training-operator.
Related Issues (20)
- Support MLX on Kubernetes with Kubeflow HOT 2
- Migrate to controller-runtime logger HOT 5
- Support CertManager for the Webhook cert generation HOT 1
- Unable to start elastic PyTorchJob example HOT 5
- Commonize webhook validations at the some points
- Update developer documentation for arm HOT 1
- Aunpun1.00 HOT 1
- Update pytorch launcher component in Kubeflow Pipelines repository HOT 3
- Update developer guide to handle missing training-operator-webhook-cert HOT 2
- Job Status is failed, when scale-in ps. HOT 4
- Failed K8s nodes leave jobs hanging indefinitely HOT 3
- Update examples for `train` API HOT 1
- [Question] Training Operator v1.8 Release Date HOT 1
- Why manifests/base/service.yaml does not include webhook server port (443) in version 1.7.0~1.5.0? HOT 7
- Not getting Kubeflow Training SDK v1.7 when installing `kubeflow-training` HOT 13
- Flaky Test: [It] should create desired Pods and Services: Distributed TFJob (4 workers, 2 PS) is succeeded
- MPIJob requires service names for the pods. HOT 3
- Add DeepSpeed Example with MPI Operator HOT 9
- chore(style): provide type for `STORAGE_INITIALIZER_VOLUME` constant
- fix(compatability): match-case syntax only compatible with Python3.10 HOT 5
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from training-operator.