- https://doi.org/10.1109/ictbig68706.2025.11324006
An Efficient FPGA System for Hypersparse Feed Forward Networks for Chip Memory Mapping
- Dec 12, 2025
- S Kumaran +5 more
For highly broad and deep models, dense floating-point computing and off-chip DRAM consumption on existing GPU and CPU-GPU platforms result in substantial index overheads (sometimes exceeding 99%) and latency bottlenecks, making it difficult to deploy hypersparse feed-forward networks (FFNs) economically. Existing techniques highlighted sparse attention aimed at combined FFN-attention optimizations; these keep index structures bloated and memory-intensive. The above shortcomings are solved by the proposed FPGA-centric system that compresses indices using hybrid CSR encodings, quantizes with 4-8 bits, and maps all data to on-chip BRAM/URAM. The proposed system prevents repeated DRAM access, diminishes index dominance, and maintains performance across multiple concurrent sparse computation engines. The proposed technique for deploying hypersparse FFNs improves CPU performance by <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$150 \times$</tex>, GPU performance by <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$6 \times$</tex>, and energy efficiency by <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$50 \times$</tex> while decreasing LUT/FF/DSP usage. However, it comes at the cost of greater BRAM consumption.