Comments (5)
Hi parth-chudasama, SDK releases > 2.14 offer support for llama2 models. Can you download one and let us know if that works for you. We plan to continue improving its accuracy and performance in subsequent releases.
from transformers-neuronx.
We are working on adding support for LLAMA in an upcoming release. Will update once we have it.
from transformers-neuronx.
Hi, I am trying to apply the current implementation of the class LlamaForSampling
. I am trying this code:
model = AutoModelForCausalLM.from_pretrained("openlm-research/open_llama_7b_700bt_preview")
save_pretrained_split(model, 'llama-split')
model_neuron = LlamaForSampling.from_pretrained('llama-split', batch_size=1, tp_degree=2, n_positions=256, amp='f32', unroll=None)
model_neuron.to_neuron()
and I am getting this error:
---------------------------------------------------------------------------
AttributeError Traceback (most recent call last)
Cell In[22], line 1
----> 1 model_neuron.to_neuron()
File ~/anaconda3/envs/llm310/lib/python3.10/site-packages/transformers_neuronx/llama/model.py:76, in LlamaForSampling.to_neuron(self)
73 new_layer.add_pre_mlp_layer_norm(layer.post_attention_layernorm.weight.detach(), None)
75 # Note: Automatic MLP padding is safe since zeros are *only* introduced to intermediary state
---> 76 new_layer.add_parameter(mlp.gate_proj.weight.T, sharding=1, allow_pad=True)
77 new_layer.add_parameter(mlp.up_proj.weight.T, sharding=1, allow_pad=True)
78 new_layer.add_parameter(mlp.down_proj.weight.T, sharding=0, allow_pad=True)
File ~/anaconda3/envs/llm310/lib/python3.10/site-packages/torch/nn/modules/module.py:1269, in Module.__getattr__(self, name)
1267 if name in modules:
1268 return modules[name]
-> 1269 raise AttributeError("'{}' object has no attribute '{}'".format(
1270 type(self).__name__, name))
AttributeError: 'DecoderLayer' object has no attribute 'add_parameter'
Do you have any ideas regarding this issue?
from transformers-neuronx.
Llama is still under dev, please follow progress here: https://github.com/aws-neuron/transformers-neuronx/tree/main/src/transformers_neuronx/llama
from transformers-neuronx.
Closing since LlamaV2 support has now been added
from transformers-neuronx.
Related Issues (20)
- Can't save/serialize any models except GPT2 HOT 3
- Compilation error on llama 7 B with batch size 8 HOT 3
- from_pretrained is broken after transformers made safetensor serialization default HOT 1
- LLaMA fails when the input token length is over 1790 tokens HOT 6
- Llama2 inference overhead time way too long HOT 6
- llama-2/codellama benchmark for inf2.xlarge HOT 4
- Mixtral Model support HOT 2
- Vicuna13B model support
- Inf2 Modified Llama 2 Loading Issue HOT 11
- Skipping generation for useless tokens, and modiying cacheids HOT 3
- How to use generate() with inputs_embeds HOT 2
- Mixtral config issue -- not handling null well HOT 8
- Generate Llama 2 from Embeddings HOT 5
- Infering logits from `model.forward` for the entire batch instead of the last forward's output. HOT 5
- Support for MPT model HOT 1
- `stopping_criteria_list(input_ids, probs)` does not check for the correct sequence. HOT 4
- User feedback when compiling and reloading a large model HOT 1
- Issue while compiling Mistral 7B 0.2 Instruct HOT 5
- Backward compatibility with saved llama 2 compiled artifacts HOT 1
- NaN outputs when masking llama model inputs HOT 6
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from transformers-neuronx.