Thanks for the repo. Have sampling working fine from your "PrettyBig" model. <p di

How to process raw text files to create similar "PrettyBig" model? about gpt2 HOT 5 CLOSED

connorjl commented on July 20, 2024

How to process raw text files to create similar "PrettyBig" model?

from gpt2.

Comments (5)

kizinfo commented on July 20, 2024

Modify create_tfrecords.py at the top, filename and path declarations point to the text files. make sure the line:
files = glob.glob(os.path.join(base_dir, "*.txt"))
properly points to and indexes the source text files.
also need a copy of the existing model encoder.
and set files_per as the number of text files to use for each chunk of tfrecords. I already had them split into chunks with ~300 books each separated by <|EndOfText|>. the .py doesn't add an end text token so if you didn't already put them in your txt files and want them need to modify the code further.

have been able to train 4gig of text on colab TPU over a number of days to 60k iterations. Results arn't as good as finetuning existing models (yet). The biggest model that still fit colab was:

"n_head": 17,
"lr": 0.00025,
"warmup_steps": 2000,
"beta1": 0.0,
"decay_exponent": 0.8,
"opt_name": "adafactor",
"decay_type": "pow",
"train_batch_size": 8,
"max_steps": 200000,
"predict_batch_size": 1,
"eval_batch_size": 8,
"iterations": 100,
"n_embd": 1020,
"n_ctx": 1024,
 "n_layer": 34,

With such small batch sizes, I'm not sure it will ever work well and without bfloat16 working on inference can't get bigger models on colab. The author says he trained models on a pod (which with 'evalutaion' pricing would cost tens of thousands of dollars).

From the original GPT2 paper the Authors claimed that the variation in training text was important for bigger models so it might be that training a model from scratch on a domain specific 4gig corpus won't ever do as well as training on a general 40gig corpus and then finetuning on the domain.

Have been able to get some really good results from the original 345M GPT2 models by finetuning on domain specific content that maintains context well through multiple paragraphs.

from gpt2.

GenTxt commented on July 20, 2024

Thanks for the information. A lot to test and as you say it's likely true that " ... a domain specific 4gig corpus won't ever do as well as training on a general 40gig corpus and then finetuning on the domain."

I'm having good results too with genre specific models based on the OpenAi 345M. Can only hope they decide to release their larger models within 6 months.

Closing this now

from gpt2.

ConnorJL commented on July 20, 2024

I'm sorry the scripts are pretty poorly documented, I'm planning on making a better custom dataset setup when I get the time. You basically just want to use create_tfrecords.py as kizinfo said to generate the .tfrecords files from your txt files.

You do NOT have to add <|endoftext|> manually! If you use my bpe_text function (in inputs.py) as input, it automatically samples "stitch" amount of texts from your dataset, concatenates them with <|endoftext|> in between and then samples n_ctx amount of tokens from the final result. Make sure that "stitch" is set so that (your minimal length text * stitch) >= n_ctx.

I plan on releasing the 1.5B model, see my blogposts about it here and here.

from gpt2.

GenTxt commented on July 20, 2024

Thanks. I got the basics working based on kizinfo's detailed reply but it didn't seem to actually reduce the error loss. I will check the scripts again as per your advice. Looking forward to testing your 1.5B model. Checking your blogs now. Cheers

…

On Thu, Jun 6, 2019 at 10:25 AM ConnorJL ***@***.***> wrote: I'm sorry the scripts are pretty poorly documented, I'm planning on making a better custom dataset setup when I get the time. You basically just want to use create_tfrecords.py as kizinfo said to generate the .tfrecords files from your txt files. You do NOT have to add <|endoftext|> manually! If you use my bpe_text function (in inputs.py) as input, it automatically samples "stitch" amount of texts from your dataset, concatenates them with <|endoftext|> in between and then samples n_ctx amount of tokens from the final result. Make sure that "stitch" is set so that (your minimal length text * stitch) >= n_ctx. I plan on releasing the 1.5B model, see my blogposts about it here ***@***.***/gpt2-counting-consciousness-and-the-curious-hacker-323c6639a3a8> and here ***@***.***/replicating-gpt2-1-5b-86454a7f26af>. — You are receiving this because you modified the open/close state. Reply to this email directly, view it on GitHub <#2?email_source=notifications&email_token=AFMAWPNSEHUT25BDOHDNDETPZEM5LA5CNFSM4HUUW4FKYY3PNVWWK3TUL52HS4DFVREXG43VMVBW63LNMVXHJKTDN5WW2ZLOORPWSZGODXDALOI#issuecomment-499516857>, or mute the thread <https://github.com/notifications/unsubscribe-auth/AFMAWPMKVACKEAQMWIA7FPLPZEM5LANCNFSM4HUUW4FA> .

from gpt2.

GenTxt commented on July 20, 2024

Enjoyed reading your blogs and I'm in full agreement. To the points you raised I'm wondering if your models can be used for fine-tuning a custom corpus model the same as the current Open-Ai 345 version? I think this is a great part of OpenAi's fear, false as it may be, that prevents them from releasing the full code. I'm getting great results fine-tuning corpus files with nshepperd's repo and selected branches using the current OpenAi 345M For example:: https://github.com/mkturkcan/GPTune (provides pre-trained models to download) based on the finetuning code released by nshepperd The one drawback is waiting for the possible release of the larger OpenAi models to fully test these repos.. It would be fantastic if your version could be tweaked to offer the same ability. Cheers, and thanks again for the great work.

…

On Thu, Jun 6, 2019 at 8:06 PM Aaron Allan ***@***.***> wrote: Thanks. I got the basics working based on kizinfo's detailed reply but it didn't seem to actually reduce the error loss. I will check the scripts again as per your advice. Looking forward to testing your 1.5B model. Checking your blogs now. Cheers On Thu, Jun 6, 2019 at 10:25 AM ConnorJL ***@***.***> wrote: > I'm sorry the scripts are pretty poorly documented, I'm planning on > making a better custom dataset setup when I get the time. You basically > just want to use create_tfrecords.py as kizinfo said to generate the > .tfrecords files from your txt files. > > You do NOT have to add <|endoftext|> manually! If you use my bpe_text > function (in inputs.py) as input, it automatically samples "stitch" amount > of texts from your dataset, concatenates them with <|endoftext|> in between > and then samples n_ctx amount of tokens from the final result. Make sure > that "stitch" is set so that (your minimal length text * stitch) >= n_ctx. > > I plan on releasing the 1.5B model, see my blogposts about it here > ***@***.***/gpt2-counting-consciousness-and-the-curious-hacker-323c6639a3a8> > and here > ***@***.***/replicating-gpt2-1-5b-86454a7f26af>. > > — > You are receiving this because you modified the open/close state. > Reply to this email directly, view it on GitHub > <#2?email_source=notifications&email_token=AFMAWPNSEHUT25BDOHDNDETPZEM5LA5CNFSM4HUUW4FKYY3PNVWWK3TUL52HS4DFVREXG43VMVBW63LNMVXHJKTDN5WW2ZLOORPWSZGODXDALOI#issuecomment-499516857>, > or mute the thread > <https://github.com/notifications/unsubscribe-auth/AFMAWPMKVACKEAQMWIA7FPLPZEM5LANCNFSM4HUUW4FA> > . >

from gpt2.

How to process raw text files to create similar "PrettyBig" model? about gpt2 HOT 5 CLOSED

Comments (5)

Related Issues (20)

Recommend Projects

React

Vue.js

Typescript

TensorFlow

Django

Laravel

D3

Recommend Topics

javascript

web

server

Machine learning

Visualization

Game

Recommend Org

Facebook

Microsoft

Google

Alibaba

D3

Tencent