Comments (4)
oh, one more piece of information, if sequence_length <=256, then I do get 35 ms/token for llama2 and codellama (unfortunately for code generation it usually needs sequence_length > 256)
from transformers-neuronx.
Thanks @zliendo . We will take a look.
from transformers-neuronx.
Thanks @zliendo I have checked the results in the blog and they are the results as of 11/7/23 when the blog was published. We are continuing to improve the performance of LLaMA on Neuron so will hope that you will try an upcoming release. Please keep an eye out for what's new and the performance page for updated performance data.
from transformers-neuronx.
Thank you so much! I will for sure try upcoming releases.
from transformers-neuronx.
Related Issues (20)
- Avoid splitting Hugging Face Hub checkpoint files on disk HOT 7
- Can't save/serialize any models except GPT2 HOT 4
- Compilation error on llama 7 B with batch size 8 HOT 4
- from_pretrained is broken after transformers made safetensor serialization default HOT 1
- LLaMA fails when the input token length is over 1790 tokens HOT 6
- Llama2 inference overhead time way too long HOT 6
- Mixtral Model support HOT 2
- Vicuna13B model support HOT 1
- Inf2 Modified Llama 2 Loading Issue HOT 11
- Skipping generation for useless tokens, and modiying cacheids HOT 3
- How to use generate() with inputs_embeds HOT 2
- Mixtral config issue -- not handling null well HOT 8
- Generate Llama 2 from Embeddings HOT 5
- Infering logits from `model.forward` for the entire batch instead of the last forward's output. HOT 6
- Support for MPT model HOT 1
- `stopping_criteria_list(input_ids, probs)` does not check for the correct sequence. HOT 4
- User feedback when compiling and reloading a large model HOT 1
- Issue while compiling Mistral 7B 0.2 Instruct HOT 5
- Backward compatibility with saved llama 2 compiled artifacts HOT 1
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from transformers-neuronx.