The rapid growth of large-scale language models and Agentic AI systems has led to a dramatic increase in data volume and memory bandwidth requirements. While High Bandwidth Memory (HBM) provides high bandwidth, its limited capacity presents challenges for next-generation AI workloads.
In this seminar, we introduce High Bandwidth Flash (HBF), a NAND-based vertically stacked memory architecture designed to deliver HBM-level bandwidth with SSD-class capacity. We present workload analyses of GPU–HBM–HBF systems for AI training and inference, and discuss hybrid memory architectures optimized for multi-modal large language models.
The seminar also covers memory-centric computing and near-memory computing architectures, along with key design considerations including signal integrity, power integrity, thermal challenges, and advanced packaging technologies such as TSVs and interposers. Based on recent research results from KAIST TERA Lab, we outline the technical roadmap for HBF and its role in future AI memory hierarchies.
ZOOM Information:
ZOOM Link:
https://us02web.zoom.us/j/2885283810?pwd=OUtNOGl0anRscURoQjRuUHkzUUFWUT09
Meeting ID: 288 528 3810
Password: kaist1234
| Time (KST) | Topic | Presenter |
|---|---|---|
| 09:00 – 10:20 | HBF Technology, Workload Analysis and Roadmap | Joungho Kim |
| 10:20 – 11:00 | Workload Analysis of GPU-HBM-HBF Architecture for Multi-Modal Large Language Model Inference | Haeseok Suh |
| 11:00 – 11:10 | Break | – |
| 11:10 – 11:40 | SI and Workload Analysis of Hybrid GDDR-HBF Memory Architecture with Multi-GPU for Efficient Disaggregated LLM Inference | Youngsoo Yoon |
| 11:40 – 12:10 | HBM–HBF Hybrid Memory Architecture for Multi-Modal AI Inference with RAG-Enhanced MoE Models | Inyoung Choi |
| 12:10 – 13:10 | Lunch | – |
| 13:10 – 13:40 | HBM–HBF Centric Memory Pooling Architecture with Custom Base Die for Transformer-Based LLM Inference | Junho Park |
| 13:40 – 14:10 | Mem-Village: A Hybrid GPU-HBM-HBF On-Glass Architecture for Ultra-Scale AI Inference Systems | Hyunyi Lee |
| 14:10 – 14:40 | HBM-HBF Based Near-Memory Computing Architecture for LLM Inference Considering Signal Integrity and Workload Analysis | Chaemin Yang |
| 14:40 – 14:50 | Break | – |
| 14:50 – 15:20 | Conditional MaskGIT-Based Imitation Learning AI Agent for P/G TSV Placement Optimization Considering HBF IR Drop | Eunji Seo |
| 15:20 – 15:50 | Thermal TSV Array Optimization Agent for Next-Generation 3D GPU–HBM Architecture Considering Thermal and Signal Integrity | Gwantak Lee |
| 15:50 – 16:20 | PDN Impedance Estimator for Multi Power Domain Interposer in Next-Generation GPU–HBM–HBF Architecture | Seungjae Lee |
Copyright ⓒ 2015 KAIST Electrical Engineering. All rights reserved. Created by PRESSCAT
Copyright ⓒ 2015 KAIST Electrical Engineering. All rights reserved. Created by PRESSCAT
Copyright ⓒ 2015 KAIST Electrical Engineering. All rights reserved. Created by PRESSCAT
Copyright ⓒ 2015 KAIST Electrical
Engineering. All rights reserved.
Created by PRESSCAT