Hey there, 😊
I want to train the DLM with a larger dataset ~17.000.000 sentences and this totally explodes my memory even though I have 200GB available. As far as I can assess, the whole training set is being tokenized in the beginning which causes the problem. Is there already a solution for it or are you aware of this problem?
This line causes the problem.
Hey there, 😊
I want to train the DLM with a larger dataset ~17.000.000 sentences and this totally explodes my memory even though I have 200GB available. As far as I can assess, the whole training set is being tokenized in the beginning which causes the problem. Is there already a solution for it or are you aware of this problem?
This line causes the problem.