굉장히 공격적이고 사나운 사람MLSys · GPU Systems · Deep Learning HardwarePart 0 EE신호와 시스템 Soliman & Srinath 교재 기준 신호와 시스템 수업 노트. 신호의 정의, 연속성과 불연속점의 값, 사각 펄스, 연속시간과 이산시간, 주기 신호와 사인파, 하모닉, 조화 관계 복소지수.디지털논리회로 Floyd 교재 기준 디지털논리회로 수업 노트. 아날로그 양과 디지털 양, 논리 준위, 펄스와 주기 파형, 클럭과 타이밍 다이어그램, NOT·AND·OR, 비교기·가산기·인코더·디코더, 레지스터와 카운터, …Part 1 Compiler01 Loop-Invariant Code Motion: From Loop Body to Preheader LLVM LICM은 반복마다 같은 곱셈과 덧셈을 반복문 본문에서 preheader로 옮긴다. 명령의 operand가 모두 반복 불변이어도 메모리 효과나 0회 실행 경로의 안전성이 증명되지 않으면 반복문 안에 남는다.00 LLVM IR and the Compilation Pipeline: From C Code to Machine Code 인터프리터, JIT, AOT의 차이에서 시작해 clang으로 C 코드를 LLVM IR로 뽑아 해독하고, llc로 어셈블리를 거쳐 실행 파일까지 내려가는 컴파일 파이프라인 전체를 다룬다. 최적화 pass가 IR을 바꾸 …Part 2 CUDA05 CUDA Concurrency: Streams, Async Copies, and Overlap 복사하는 동안 계산도 같이 하면 좋겠죠ㅎㅎ host memory와 device memory부터 pinned memory, cudaMemcpyAsync, stream, chunk까지 살펴보며 데이터 복사와 kernel …04 CUDA Unified Memory: Virtual Address, Placement, and Coherence CPU와 GPU가 메모리 하나를 함께 쓰면 안에서는 무슨 일이 생길까요? 가상 주소와 데이터 배치·이동, 동기화와 캐시 일관성을 살펴보고 Jetson AGX Orin의 실제 장치 속성에 대입해봐요.03 CUDA Shared Memory: Tiling, Bank Conflicts, and Reduction global memory를 읽어 shared memory에 재사용할 때 무엇을 챙겨야 할까요? Coalescing과 tiling부터 bank conflict를 푸는 padding·swizzle, occupancy, …02 CUDA C Basics CUDA C에서 입력을 GPU에 보내고 결과를 받아오는 과정을 따라가봐요. Host-device memory와 kernel launch, thread·block·warp의 배치를 살펴보고, occupancy …01 NVIDIA GPU Architecture Genealogy: Tesla to Rubin Tesla(2006)부터 Rubin(2026)까지 NVIDIA GPU의 계보를 따라가요. SIMT·warp=32·SM당 block은 유지하면서 어떤 전용 연산기를 더했는지, Tensor Core …00 GPU Architecture Primer: The Tesla Foundation CUDA C에 나오는 warp, SM, coalescing은 어디서 왔을까요? 2008 IEEE Micro 논문의 Tesla(G80/GT200)를 따라 Input Assembler …Part 3 MIT 6.59306.5930 L02 - From Einsum to DNN Workloads L02 전체: accelerator 설계 절차, tensor와 Einsum, iteration space, memory traffic, compute intensity, Roofline, CNN …6.5930 L01 - Introduction and Applications 딥러닝 하드웨어가 왜 따로 필요한가. 비싼 data movement, 범용 CPU의 비용, domain-specific accelerator와 TeAAL, Roofline까지 L01의 논리를 한 줄로 잇는다.MIT 6.5930: Hardware Architecture for Deep Learning MIT EECS 6.5930 by Prof. Vivienne Sze and Prof. Joel Emer, Hardware Architecture for Deep LearningPart 4 EmbeddedJetPack 6.2.2 Flash Troubleshooting: AMD USB 비호환과 chroot 해법 Jetson AGX Orin을 플래시할 때 AMD 호스트에서 발생하는 tegrarcm_v2 USB write timeout의 근본 원인, 그리고 RAM이 부족한 Intel 노트북을 사용한 chroot 기반 우회법을 …Part 5 MIT 6.59406.5940 L01: Introduction and Overview MIT 6.5940 (Song Han) Lecture 1 정리. DNN 효율화의 필요성부터 Model Compression, Quantization(AWQ/SmoothQuant/RTN), Sparsity, Edge …Part 6 GeneralOpening Note GPU Systems & Deep Learning Hardware