Comments (1)
Hi @zhaotyer could you provide some additional information about how you collected these numbers? Are you running the benchmark in our DeepSpeedExamples repo?
If so, are you gathering these numbers directly from the resulting log files?
I just ran the Llama-2-7b model on 1xA6000 GPU with prompt size 256 and generation size 256 for 1, 2, 4, 8, 16, and 32 clients and I'm seeing roughly equal performance for vLLM and FastGen (DeepSpeed-MII):
This is expected for the current release. FastGen is capable of providing better performance with longer prompts and shorter generation lengths. We go into greater detail of the performance and benchmarks in the two FastGen release blogs here and here.
from deepspeed-mii.
Related Issues (20)
- support stream
- [BUG] MII Backend Hangs After 9999 Exceptions in `MIIAsyncPipeline.put_request` HOT 2
- few questions regarding the implementation of streaming and batching
- Configure server log level HOT 2
- Compute perplexity
- Attempting to flush sequence N which does not exist
- deepseed-mii支持多节点推理么 HOT 2
- Import Error, not compatible with transformer package HOT 4
- CUDA device rank in mii.pipeline
- Client cannot find deployment error
- Dummy data loading?
- non-persistent simple example does not work HOT 5
- non-persistent example doesn't work on Mixtral-8*7B-v0.1
- Can't use Llama 3.1 with MII, ImportError: cannot import name 'Conversation' from 'transformers' HOT 1
- FileExistsError: [Errno 17] File exists: '/tmp/mii_cache' ` on generate function call
- By default does deepspeed mii use bf16 dtype or fp16?
- OpenAI server fails
- Configuration setting to pass parameters to tokenizer while encoding and decoding
- Question About Offloading and Recomputation
- multi model deployment HOT 1
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from deepspeed-mii.