Developer benchmarks GPU generations for local LLM speed in nightly stock analysis pipeline
A developer running a nightly stock-analysis AI pipeline has published detailed benchmarks comparing local LLM inference speeds across multiple GPU generations, from dual RTX 4070 Ti Super cards to the RTX 5090. The production system currently uses a Qwen3-35B MoE model quantized to Q4_K_M, processing around 55 million input tokens and 7.5 million output tokens per weekday across roughly 5,200 API calls. On the same RTX 5090 hardware, the MoE model completed a 100-stock batch in 46 minutes at 911 tokens per second, while a 27B dense model took 5 hours 41 minutes at 124 tokens per second — roughly 7.3 times slower. The developer attributes the gap to fundamental differences in model architecture rather than hardware limitations, noting that MoE activates far fewer parameters per token than a dense model. A final decision on which model to adopt in production is pending further evaluation, including self-consistency scoring and forward rank correlation against actual stock returns.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in