The spec dec performance of Eagleis worse than expected as shown below:
Model: meta-llama/Meta-Llama-3.1-70B-Instruct
Draft model: yuhuili/EAGLE-LLaMA3-Instruct-70B
Hardware: 4xH100
Target model TP=4
Dataset: ShareGPT
vllm version: v0.6.1.post2
Even at low QPS, the performance is far from 2x speedup reported in the original eagle paper (light blue line is the original, the solid lines are with SD). We need to understand the performance gap here. Possible reasons include but not limited to
Profiling is required to understand the issue. Open this issue to track the progress.
Report of performance regressionNo response
Misc discussion on performanceNo response
Your current environment (if you think it is necessary)The output of `python collect_env.py`
Before submitting a new issue...
RetroSearch is an open source project built by @garambo | Open a GitHub Issue
Search and Browse the WWW like it's 1997 | Search results from DuckDuckGo
HTML:
3.2
| Encoding:
UTF-8
| Version:
0.7.4