News & Event​

HBF Technology: Workload Analysis and Roadmap

Subject

HBF Technology: Workload Analysis and Roadmap

Date

2026. 02. 10.(Tue) 10am~

Speaker

Professor Joungho Kim (KAIST) and KAIST TERA Lab Researchers

Place

Online Seminar (ZOOM)

Overview:

The rapid growth of large-scale language models and Agentic AI systems has led to a dramatic increase in data volume and memory bandwidth requirements. While High Bandwidth Memory (HBM) provides high bandwidth, its limited capacity presents challenges for next-generation AI workloads.

In this seminar, we introduce High Bandwidth Flash (HBF), a NAND-based vertically stacked memory architecture designed to deliver HBM-level bandwidth with SSD-class capacity. We present workload analyses of GPU–HBM–HBF systems for AI training and inference, and discuss hybrid memory architectures optimized for multi-modal large language models.

The seminar also covers memory-centric computing and near-memory computing architectures, along with key design considerations including signal integrity, power integrity, thermal challenges, and advanced packaging technologies such as TSVs and interposers. Based on recent research results from KAIST TERA Lab, we outline the technical roadmap for HBF and its role in future AI memory hierarchies.

ZOOM Information:

ZOOM Link:
https://us02web.zoom.us/j/2885283810?pwd=OUtNOGl0anRscURoQjRuUHkzUUFWUT09

Meeting ID: 288 528 3810

Password: kaist1234

[Program Schedule]

Time (KST) Topic Presenter
09:00 – 10:20 HBF Technology, Workload Analysis and Roadmap Joungho Kim
10:20 – 11:00 Workload Analysis of GPU-HBM-HBF Architecture for Multi-Modal Large Language Model Inference Haeseok Suh
11:00 – 11:10 Break
11:10 – 11:40 SI and Workload Analysis of Hybrid GDDR-HBF Memory Architecture with Multi-GPU for Efficient Disaggregated LLM Inference Youngsoo Yoon
11:40 – 12:10 HBM–HBF Hybrid Memory Architecture for Multi-Modal AI Inference with RAG-Enhanced MoE Models Inyoung Choi
12:10 – 13:10 Lunch
13:10 – 13:40 HBM–HBF Centric Memory Pooling Architecture with Custom Base Die for Transformer-Based LLM Inference Junho Park
13:40 – 14:10 Mem-Village: A Hybrid GPU-HBM-HBF On-Glass Architecture for Ultra-Scale AI Inference Systems Hyunyi Lee
14:10 – 14:40 HBM-HBF Based Near-Memory Computing Architecture for LLM Inference Considering Signal Integrity and Workload Analysis Chaemin Yang
14:40 – 14:50 Break
14:50 – 15:20 Conditional MaskGIT-Based Imitation Learning AI Agent for P/G TSV Placement Optimization Considering HBF IR Drop Eunji Seo
15:20 – 15:50 Thermal TSV Array Optimization Agent for Next-Generation 3D GPU–HBM Architecture Considering Thermal and Signal Integrity Gwantak Lee
15:50 – 16:20 PDN Impedance Estimator for Multi Power Domain Interposer in Next-Generation GPU–HBM–HBF Architecture Seungjae Lee

Profile: