News
Newest
Ask
Show
Jobs
Open on GitHub
Show HN: A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo)
(mikeayles.com)
7 points | by
mikeayles
4 hours ago
3 comments
haeseong
2 hours ago
I didn't expect the 2,000 connection sweep to stay flat, since all of them are sharing one stream. What does per user latency look like at that end of the sweep?
[-]
mikeayles
2 hours ago
[flagged]
mikeayles
4 hours ago
[dead]
threadsnoop
2 hours ago
[dead]
3 comments