Currently, PatrickStart could train the largest pretrained model with the lowest hardware requirement, i.e. gpu and cpu memory. However, this comes with a price, as we are specializing the design to naive bert and gpt structure as well as adam optimizer. This makes our users hard to use PatrickStar for their latest research project and hard for us to tweak some edge cases to be compatible with popular NLP repos.
Therefore, we decide to refactor PatrickStar to make it simple and flexible. After all, comparing to break record, we prefer creating a handy tool to the NLP community :)
Here are some of the changes we are making now:
We will try to make the new design as performant and as efficient as the old one, however if what you need is the extreme performance mentioned as the paper, please refer to release v0.4.6.
Currently, PatrickStart could train the largest pretrained model with the lowest hardware requirement, i.e. gpu and cpu memory. However, this comes with a price, as we are specializing the design to naive bert and gpt structure as well as adam optimizer. This makes our users hard to use PatrickStar for their latest research project and hard for us to tweak some edge cases to be compatible with popular NLP repos.
Therefore, we decide to refactor PatrickStar to make it simple and flexible. After all, comparing to break record, we prefer creating a handy tool to the NLP community :)
Here are some of the changes we are making now:
We will try to make the new design as performant and as efficient as the old one, however if what you need is the extreme performance mentioned as the paper, please refer to release v0.4.6.