CURRICULUM VITAE
Jie Ren
Computer Science · KAUST
A print-friendly version: use your browser’s Print → Save as PDF.
Research Interests
My research focuses on the co-design of numerical algorithms and modern heterogeneous architectures. I study sparse factorizations, finite element methods, GPU scheduling, and architecture-aware kernel design, with an emphasis on reformulating classical numerical methods for extreme parallelism.
Education
- Ph.D. in CS, KAUST, 2023 -- Present
- Advisor: Prof. David E. Keyes
- B.Eng. in Electronics Engineering, Xidian University, 2017 -- 2021
Technical Skills
- Programming: C++ (17/20), CUDA, Python, MATLAB
- GPU Systems: Architecture-aware algorithm design, memory hierarchy optimization, instruction-level parallelism, kernel fusion, and fine-grained task scheduling
- Frameworks & Libraries: CUTLASS, libCEED, MFEM, Eigen, PyTorch, and NumPy
Experiences
May 2025 -- August 2025
- Baidu, Beijing, China
- Mentor: Prof. Haifeng Wang, CTO
- Role: ERNIE Star Research Intern
- Developed optimized INT2 GPU kernels based on Stream-K, achieving a 5x speedup over the Triton baseline. Redesigned a high-speed-train CFD pipeline and implemented custom GPU kernels, achieving a 3x speedup while reducing memory use by 50%.
July 2024 -- October 2024
- NVIDIA Research (Networking Group), Santa Clara, CA, USA
- Mentor: Dr. Nic McDonald
- Role: Research Intern
- Designed communication-computation-overlapped tensor-parallel GEMM kernels for distributed inference. The work was upstreamed into NVIDIA CUTLASS.
February 2023 -- October 2023
- ETH Zurich, Switzerland
- Host: Prof. Torsten Hoefler
- Role: Academic Guest (funded by HiPEAC Grant)
- Contributed to INT4 Transformer quantization with CUTLASS, achieving a 4x speedup, and developed mixed-precision GPU-centric sparse linear solvers.
July 2021 -- September 2023
- The University of Edinburgh, UK
- Supervisor: Prof. Luo Mai
- Role: Academic Guest
- Led the design and implementation of TorchOpt, a high-performance differentiable optimization library published in JMLR, and contributed to mixed-precision sparse linear solver research.
February 2022 -- October 2022
- MegEngine, MEGVII Inc., Beijing, China
- Mentor: Biao Wang
- Role: Software Engineering Intern
- Migrated depth-wise convolution kernels from Turing to Ampere by rewriting the CUTLASS global-memory iterator, achieving a 5x+ speedup, and implemented a region-restricted convolution operator.
April 2021 -- February 2022
- 3D Vision, MEGVII Inc., Beijing, China
- Mentor: Ran Yan
- Role: Research Intern
- Designed MegBA, a distributed GPU library for large-scale bundle adjustment published at ECCV 2022. Reduced HF-Net inference latency from 20 ms to 5 ms on an RTX 2080 Ti.
Publications
- J. Ren, H. Ltaief, S. Zampini, and D. E. Keyes. “Cheetah: Optimizing Execution Pipelines for Matrix-Free Finite Element Operators on GPUs.”
ACM International Conference on Supercomputing (ICS), 2026. - L. Wang, J. Wang, J. Ren, Z. Xiang, D. E. Keyes, and D. Wang. “Private Training of Large-Scale Models with Efficient DP-SGD.”
Conference on Neural Information Processing Systems (NeurIPS), 2025. - L. Wang, J. Ren, H. Xu, J. Wang, D. E. Keyes, and D. Wang. “ZO-Offloading: Fine-Tuning LLMs with 100 Billion Parameters on a Single GPU.”
Conference on Language Modeling (COLM), 2025. - J. Ren, T. Zhong, Y. Hong, G. Feng, X. Wang, W. Jia, H. Ltaief, and D. E. Keyes. “Caracal: A GPU-Resident Sparse LU Solver with Lightweight Fine-Grained Scheduling.”
International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), 2025. - J. Ren, H. Ltaief, S. Abdullah, and D. E. Keyes. “Accelerating Mixed-Precision Out-of-Core Cholesky Factorization with Static Task Scheduling.”
ISC High Performance, 2025. - S. Ashkboos, I. Markov, E. Frantar, T. Zhong, X. Wang, J. Ren, T. Hoefler, and D. Alistarh. “Towards End-to-End 4-Bit Inference on Generative Large Language Models.”
Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024. - H. Ltaief, R. Alomairy, Q. Cao, J. Ren, L. Slim, T. Kurth, B. Dorschner, S. Bougouffa, R. Abdelkhalak, and D. E. Keyes. “Toward Capturing Genetic Epistasis from Multivariate Genome-Wide Association Studies Using Mixed-Precision Kernel Ridge Regression.”
International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), 2024. Gordon Bell Prize Finalist. - J. Ren, B. Liu, X. Feng, X. Pan, Y. Fu, Y. Yang, and L. Mai. “TorchOpt: An Efficient Library for Differentiable Optimization.”
Journal of Machine Learning Research (JMLR), 2023. - B. Liu, X. Feng, J. Ren, L. Mai, R. Zhu, H. Zhang, J. Wang, and Y. Yang. “A Theoretical Understanding of Gradient Bias in Meta-Reinforcement Learning.”
Conference on Neural Information Processing Systems (NeurIPS), 2022. - X. Xia, W. Yang, J. Ren, Y. Li, Y. Zhan, B. Han, and T. Liu. “Pluralistic Image Completion with Probabilistic Mixture-of-Experts.”
Conference on Neural Information Processing Systems (NeurIPS), 2022. - J. Ren, W. Liang, R. Yan, L. Mai, S. Liu, and X. Liu. “MegBA: A High-Performance and Distributed Library for Large-Scale Bundle Adjustment.”
European Conference on Computer Vision (ECCV), 2022. - L. Tian, B. Chen, J. Ren, H. Zhang, Z. Wu, N. Han, Y. Chen, and H. Liu. “Multi-Scale Visual Attention for Attribute Disambiguation in Zero-Shot Learning.”
Signal Processing: Image Communication, 2021. - Z. Duan, D. Wang, B. Chen, C. Wang, W. Chen, Y. Li, J. Ren, and M. Zhou. “Sawtooth Factorial Topic Embeddings Guided Gamma Belief Network.”
International Conference on Machine Learning (ICML), 2021.
Preprints
- J. Ren, Y. Li, Z. Ding, W. Pan, and H. Dong. “Probabilistic Mixture-of-Experts for Efficient Deep Reinforcement Learning.”
arXiv:2104.09122, 2021.
Selected Academic Activities
- Invited Participant, Dagstuhl Seminar 26392 (September 2026)
- Young Researcher, 13th Heidelberg Laureate Forum (September 2026)
Service
Conference Reviewer
- International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), 2026
- IEEE International Conference on Robotics and Automation (ICRA), 2025
- International Conference on Parallel Processing (ICPP), 2024
- International Conference on Learning Representations (ICLR), 2023–2025
- Conference on Neural Information Processing Systems (NeurIPS), 2022–2023
- International Conference on Machine Learning (ICML), 2022–2024
- IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2021–2022
Journal Reviewer
- ACM Transactions on Mathematical Software (ACM TOMS)
- Parallel Computing
- International Journal of Computer Vision (IJCV)
- IEEE Robotics and Automation Letters (IEEE RA-L)
Awards
- ACM/IEEE-CS George Michael Memorial High Performance Computing Fellowship (2026)
- Gordon Bell Prize Finalist, SC (2024)
- KAUST Dean's Award (2024)
- 1st Place, ISC Student Cluster Competition Coding Challenge (2024)
- HiPEAC Travel & Collaboration Grant (2023)