About the journal
Browse
Collections
Multimedia collections
Authors & reviewers
DFU-L: A Latency-Sensitive Dataflow Architecture with NoC and Streaming Optimizations for Edge DSP
Haibin Wu , Zhihua Fan , Zhen Wang , Zirui Ma , Wenxing Li , Haoran Tong , Tengfei Xia , Xiaochun Ye , Wenming Li
Dataflowarchitectures are attractive for real-time edgeworkloads because structured computation and streaming data movement expose abundant instruction-level and pipeline parallelism. However, as working sets shrink and iteration latency becomes critical, conventional spatial dataflow fabrics can become limited by on-chip transfer and control overheads rather than raw compute throughput. Fine-grained inter-PE dependences, scratchpad-bank contention, and frequent instruction or kernel switching introduce non-negligible stalls, reducing effective PE utilization. This paper presents DFU-L, a latency-sensitive coarse-grained dataflow architecture that improves latency overlap across on-chip communication, streaming memory access, and execution switching. DFU-L first introduces a forwarding NoC with dataflow-unblocking arbitration to accelerate inter-PE operand delivery and reduce dependence-transfer latency. It then organizes memory accesses into clustered fetch domains and uses a paced streaming engine to smooth scratchpad-bank service under memory-bound thin-DFG execution. Finally, DFU-L reuses idle forwarding windows to overlap instruction and kernel switching with the tail of the current dataflow execution, reducing control-transition overhead. We evaluate DFU-L with a cycle-accurate simulator calibrated against RTL and a 12 nm, 1 GHz implementation. Across representative edge DSP kernels, DFU-L achieves up to 1.92× speedup and 1.45× on average over a mesh baseline. It improves energy efficiency by 1.14×–2.06× over Plasticine, reduces energy-delay product by up to 2.18× compared with state-of-the-art DSP accelerators, and achieves 1.11–1.37× higher normalized energy efficiency than advanced dataflow designs, under analytical process-technology normalization.
Dataflow architecture / edge DSP workloads / network-on-chip / streaming memory access / latency-sensitive execution
Higher Education Press 2026
/
| 〈 |
|
〉 |