Comments (3)
Hi @zsaladin - the cuda runtime version and the cuda version that is bundled with torch can be different, so that is the reason why we need to be able to check the actual cuda runtime version that is installed.
from deepspeed.
@loadams Thanks for replying. I have some questions about your answer.
-
As you mentioned versions of cuda runtime installed globally and bundled with torch can be different. The cuda runtime bundled with torch is in virtual environment.
Does deepspeed use cuda runtime in virtual environment? If deepspeed uses cuda runtime in virtual environment then the version conflict cannot happen. So it would be great that deepspeed will use bundled cuda runtime if deepspeed doesn't use it now. -
When I use
deepspeed==0.12.6
it doesn't requirenvcc
.nvcc
is a compiler not for checking version.
I'm not sure that but deepspeed needs to complie cuda code?
from deepspeed.
Hi @zsaladin - DeepSpeed uses the version of cuda runtime that is installed on the system, it cannot "use" the version that torch is built with, as that doesn't have nvcc
/cuda drivers, it is just what the installed pytorch is built against.
As for why it didn't require nvcc in 0.12.6, we will have to check the code to see what changes have taken place that would cause this. With 0.12.6 does DeepSpeed detect that you are using an Nvidia GPU? Are you able to run nvidia-smi on your system with DeepSpeed 0.14.x?
from deepspeed.
Related Issues (20)
- [BUG]I found that the parameters of model will be fully transferred to the VRAM of each process. Is this abnormal in my understanding? HOT 5
- [BUG] fp6 canβt load qwen1.5-34b-chat
- [BUG] deepspeed amp seems to convert all input to specific dtype
- Data Loading for DeepSpeed Ulysses and Data Parallelism
- different setting for same (num_gpus * batch_size * grad_accum_steps) output different loss and gradient norm HOT 1
- [BUG] Stage 3 in WSL2 throws RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu! HOT 3
- [BUG] DeepSpeed is loads the whole model to every GPUs instead of partitioning HOT 1
- [BUG] RepeatingLoader may be invalid in the pipe stages neither the fist nor last
- [REQUEST] Supporting custom generation loop (outlines, LMQL, guidance) in DeepSpeedHybridEngine
- [BUG] M1 Mac has an issue with `hostname -I` not being a valid command HOT 2
- [BUG] CUDA OOM error when Hugging Face `ignore_mismatched_sizes` is enabled
- [BUG] Zero3 causes AttributeError: 'NoneType' object has no attribute 'numel' in continual training HOT 2
- [BUG] cannot import name '_get_socket_with_port' from 'torch.distributed.elastic.agent.server.api' HOT 3
- # [REQUEST] Upstream modifications of PaRO
- Reset Optimizer HOT 1
- nv-ds-chat CI test failure
- [HELP] ZeRO3 partition parameters after fully load to each GPU! HOT 3
- [BUG] ZeRO optimizer with MoE Expert Parallelism HOT 1
- [BUG] Pipeline Dataloader Samler: `shuffle=False`
- [REQUEST] Moving a trainable model with an optimiser between GPU and CPU
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
π Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. πππ
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google β€οΈ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from deepspeed.