DFU-L: A Latency-Sensitive Dataflow Architecture with NoC and Streaming Optimizations for Edge DSP

Haibin Wu , Zhihua Fan , Zhen Wang , Zirui Ma , Wenxing Li , Haoran Tong , Tengfei Xia , Xiaochun Ye , Wenming Li

Front. Comput. Sci. ››

PDF (9521KB)
Front. Comput. Sci. ›› DOI: 10.1007/s11704-026-61054-2
RESEARCH ARTICLE
DFU-L: A Latency-Sensitive Dataflow Architecture with NoC and Streaming Optimizations for Edge DSP
Author information +
History +
PDF (9521KB)

Abstract

Dataflowarchitectures are attractive for real-time edgeworkloads because structured computation and streaming data movement expose abundant instruction-level and pipeline parallelism. However, as working sets shrink and iteration latency becomes critical, conventional spatial dataflow fabrics can become limited by on-chip transfer and control overheads rather than raw compute throughput. Fine-grained inter-PE dependences, scratchpad-bank contention, and frequent instruction or kernel switching introduce non-negligible stalls, reducing effective PE utilization. This paper presents DFU-L, a latency-sensitive coarse-grained dataflow architecture that improves latency overlap across on-chip communication, streaming memory access, and execution switching. DFU-L first introduces a forwarding NoC with dataflow-unblocking arbitration to accelerate inter-PE operand delivery and reduce dependence-transfer latency. It then organizes memory accesses into clustered fetch domains and uses a paced streaming engine to smooth scratchpad-bank service under memory-bound thin-DFG execution. Finally, DFU-L reuses idle forwarding windows to overlap instruction and kernel switching with the tail of the current dataflow execution, reducing control-transition overhead. We evaluate DFU-L with a cycle-accurate simulator calibrated against RTL and a 12 nm, 1 GHz implementation. Across representative edge DSP kernels, DFU-L achieves up to 1.92× speedup and 1.45× on average over a mesh baseline. It improves energy efficiency by 1.14×–2.06× over Plasticine, reduces energy-delay product by up to 2.18× compared with state-of-the-art DSP accelerators, and achieves 1.11–1.37× higher normalized energy efficiency than advanced dataflow designs, under analytical process-technology normalization.

Keywords

Dataflow architecture / edge DSP workloads / network-on-chip / streaming memory access / latency-sensitive execution

Cite this article

Download citation ▾
Haibin Wu, Zhihua Fan, Zhen Wang, Zirui Ma, Wenxing Li, Haoran Tong, Tengfei Xia, Xiaochun Ye, Wenming Li. DFU-L: A Latency-Sensitive Dataflow Architecture with NoC and Streaming Optimizations for Edge DSP. Front. Comput. Sci. DOI:10.1007/s11704-026-61054-2

登录浏览全文

4963

注册一个新账户 忘记密码

References

Rights & permissions

Higher Education Press 2026

PDF (9521KB)

0

Accesses

0

Citation

Detail

Sections
Recommended

/