Tuesday, 20 August 2019

Why is the memory in GPU still in use after clearing the object?

Starting with zero usage:

>>> import gc
>>> import GPUtil
>>> import torch
>>> GPUtil.showUtilization()
| ID | GPU | MEM |
------------------
|  0 |  0% |  0% |
|  1 |  0% |  0% |
|  2 |  0% |  0% |
|  3 |  0% |  0% |

Then I create a big enough tensor and hog the memory:

>>> x = torch.rand(10000,300,200).cuda()
>>> GPUtil.showUtilization()
| ID | GPU | MEM |
------------------
|  0 |  0% | 26% |
|  1 |  0% |  0% |
|  2 |  0% |  0% |
|  3 |  0% |  0% |

Then I tried several ways to see if the tensor disappears.

Attempt 1: Detach, send to CPU and overwrite the variable

No, doesn't work.

>>> x = x.detach().cpu()
>>> GPUtil.showUtilization()
| ID | GPU | MEM |
------------------
|  0 |  0% | 26% |
|  1 |  0% |  0% |
|  2 |  0% |  0% |
|  3 |  0% |  0% |

Attempt 2: Delete the variable

No, this doesn't work either

>>> del x
>>> GPUtil.showUtilization()
| ID | GPU | MEM |
------------------
|  0 |  0% | 26% |
|  1 |  0% |  0% |
|  2 |  0% |  0% |
|  3 |  0% |  0% |

Attempt 3: Use the torch.cuda.empty_cache() function

Seems to work, but it seems that there are some lingering overheads...

>>> torch.cuda.empty_cache()
>>> GPUtil.showUtilization()
| ID | GPU | MEM |
------------------
|  0 |  0% |  5% |
|  1 |  0% |  0% |
|  2 |  0% |  0% |
|  3 |  0% |  0% |

Attempt 4: Maybe clear the garbage collector.

No, 5% is still being hogged

>>> gc.collect()
0
>>> GPUtil.showUtilization()
| ID | GPU | MEM |
------------------
|  0 |  0% |  5% |
|  1 |  0% |  0% |
|  2 |  0% |  0% |
|  3 |  0% |  0% |

Attempt 5: Try deleting torch altogether (as if that would work when del x didn't work -_- )

No, it doesn't...*

>>> del torch
>>> GPUtil.showUtilization()
| ID | GPU | MEM |
------------------
|  0 |  0% |  5% |
|  1 |  0% |  0% |
|  2 |  0% |  0% |
|  3 |  0% |  0% |

And then I tried to check gc.get_objects() and it looks like there's still quite a lot of odd THCTensor stuff inside...

Any idea why is the memory still in use after clearing the cache?



from Why is the memory in GPU still in use after clearing the object?

No comments:

Post a Comment