Currently, only LLaMA-7B is supported since I haven't figured out how to

Leave a tip jar to get a <a class="user-mention notranslate" data-hovercard-type="user

Fixed with <a class="commit-link" data-hovercard-type="commit" data-hovercard-url="htt

Merging tensors of larger models about llama.cpp HOT 4 CLOSED

kir-gadjello commented on May 18, 2024 2

Merging tensors of larger models

from llama.cpp.

Comments (4)

ggerganov commented on May 18, 2024 5

Thanks! The bigger problem now is that I am out of disk space, haha!
Anyway, will try to figure out something later

from llama.cpp.

theontho commented on May 18, 2024 2

Leave a tip jar to get a @ggerganov bigger SSD and / or macbook :D

from llama.cpp.

eous commented on May 18, 2024

Its kinda pointless now but I was able to merge the 30B and 65B with this core bit of hackery added to the convert script.

+    fname_model = sys.argv[1] + "/consolidated." + str(i).zfill(2) + ".pth"
+    model_i = torch.load(fname_model, map_location="cpu")
+    
+    # Since the models are split, we need to append the tensors changing the shape/size
+    for k, v in model_i.items():
+        if k in model:
+            if model[k].dtype != v.dtype:
+                print("ERROR: Tensor types do not match: ", model[k].dtype, " vs ", v.dtype)
+                sys.exit(1)
+            elif len(model[k].shape) == 1:
+                print("Skipping tensor: " + k + " with shape: ", v.shape, " and type: ", v.dtype)
+                continue
+            elif k == "output.weight":
+                print("Concatenating tensor: " + k + " with shape: ", v.shape, " and type: ", v.dtype)
+                model[k] = torch.cat((model[k], v), dim=0)
+                print("New shape: ", model[k].shape)                
+                continue
+            elif "tok_embeddings" in k:
+                print("Concatenating tensor: " + k + " with shape: ", v.shape, " and type: ", v.dtype)
+                model[k] = torch.cat((model[k], v), dim=1)
+                print("New shape: ", model[k].shape)
+                continue
+            elif "attention.wo" in k:
+                print("Concatenating tensor: " + k + " with shape: ", v.shape, " and type: ", v.dtype)
+                model[k] = torch.cat((model[k], v), dim=1)
+                print("New shape: ", model[k].shape)
+                continue
+            elif "feed_forward.w2" in k:
+                print("Concatenating tensor: " + k + " with shape: ", v.shape, " and type: ", v.dtype)
+                model[k] = torch.cat((model[k], v), dim=1)
+                print("New shape: ", model[k].shape)
+            else:
+                print("Concatenating tensor: " + k + " with shape: ", v.shape, " and type: ", v.dtype, " with shape: ", model[k].shape)
+                model[k] = torch.cat((model[k], v), dim=0)
+                print("New shape: ", model[k].shape)
+        else:
+            print("Adding tensor: " + k + " with shape: ", v.shape, " and type: ", v.dtype)
+            model[k] = v
+    del model_i```

from llama.cpp.

ggerganov commented on May 18, 2024

Fixed with 007a8f6

On startup, we go through all the parts and merge them dynamically in the ggml buffers.

from llama.cpp.

Recommend Projects