jie.ren

CURRICULUM VITAE

Jie Ren

Computer Science · KAUST

Research Interests

My research focuses on the co-design of numerical algorithms and modern heterogeneous architectures. I study sparse factorizations, finite element methods, GPU scheduling, and architecture-aware kernel design, with an emphasis on reformulating classical numerical methods for extreme parallelism.

Education

  • Ph.D. in CS, KAUST, 2023 -- Present
  • B.Eng. in Electronics Engineering, Xidian University, 2017 -- 2021

Technical Skills

  • Programming: C++ (17/20), CUDA, Python, MATLAB
  • GPU Systems: Architecture-aware algorithm design, memory hierarchy optimization, instruction-level parallelism, kernel fusion, and fine-grained task scheduling
  • Frameworks & Libraries: CUTLASS, libCEED, MFEM, Eigen, PyTorch, and NumPy

Experiences

  • May 2025 -- August 2025

    • Baidu, Beijing, China
    • Mentor: Prof. Haifeng Wang, CTO
    • Role: ERNIE Star Research Intern
    • Developed optimized INT2 GPU kernels based on Stream-K, achieving a 5x speedup over the Triton baseline. Redesigned a high-speed-train CFD pipeline and implemented custom GPU kernels, achieving a 3x speedup while reducing memory use by 50%.
  • July 2024 -- October 2024

    • NVIDIA Research (Networking Group), Santa Clara, CA, USA
    • Mentor: Dr. Nic McDonald
    • Role: Research Intern
    • Designed communication-computation-overlapped tensor-parallel GEMM kernels for distributed inference. The work was upstreamed into NVIDIA CUTLASS.
  • February 2023 -- October 2023

    • ETH Zurich, Switzerland
    • Host: Prof. Torsten Hoefler
    • Role: Academic Guest (funded by HiPEAC Grant)
    • Contributed to INT4 Transformer quantization with CUTLASS, achieving a 4x speedup, and developed mixed-precision GPU-centric sparse linear solvers.
  • July 2021 -- September 2023

    • The University of Edinburgh, UK
    • Supervisor: Prof. Luo Mai
    • Role: Academic Guest
    • Led the design and implementation of TorchOpt, a high-performance differentiable optimization library published in JMLR, and contributed to mixed-precision sparse linear solver research.
  • February 2022 -- October 2022

    • MegEngine, MEGVII Inc., Beijing, China
    • Mentor: Biao Wang
    • Role: Software Engineering Intern
    • Migrated depth-wise convolution kernels from Turing to Ampere by rewriting the CUTLASS global-memory iterator, achieving a 5x+ speedup, and implemented a region-restricted convolution operator.
  • April 2021 -- February 2022

    • 3D Vision, MEGVII Inc., Beijing, China
    • Mentor: Ran Yan
    • Role: Research Intern
    • Designed MegBA, a distributed GPU library for large-scale bundle adjustment published at ECCV 2022. Reduced HF-Net inference latency from 20 ms to 5 ms on an RTX 2080 Ti.

Publications

  1. J. Ren, H. Ltaief, S. Zampini, and D. E. Keyes. “Cheetah: Optimizing Execution Pipelines for Matrix-Free Finite Element Operators on GPUs.”
    ACM International Conference on Supercomputing (ICS), 2026.
  2. L. Wang, J. Wang, J. Ren, Z. Xiang, D. E. Keyes, and D. Wang. “Private Training of Large-Scale Models with Efficient DP-SGD.”
    Conference on Neural Information Processing Systems (NeurIPS), 2025.
  3. L. Wang, J. Ren, H. Xu, J. Wang, D. E. Keyes, and D. Wang. “ZO-Offloading: Fine-Tuning LLMs with 100 Billion Parameters on a Single GPU.”
    Conference on Language Modeling (COLM), 2025.
  4. J. Ren, T. Zhong, Y. Hong, G. Feng, X. Wang, W. Jia, H. Ltaief, and D. E. Keyes. “Caracal: A GPU-Resident Sparse LU Solver with Lightweight Fine-Grained Scheduling.”
    International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), 2025.
  5. J. Ren, H. Ltaief, S. Abdullah, and D. E. Keyes. “Accelerating Mixed-Precision Out-of-Core Cholesky Factorization with Static Task Scheduling.”
    ISC High Performance, 2025.
  6. S. Ashkboos, I. Markov, E. Frantar, T. Zhong, X. Wang, J. Ren, T. Hoefler, and D. Alistarh. “Towards End-to-End 4-Bit Inference on Generative Large Language Models.”
    Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024.
  7. H. Ltaief, R. Alomairy, Q. Cao, J. Ren, L. Slim, T. Kurth, B. Dorschner, S. Bougouffa, R. Abdelkhalak, and D. E. Keyes. “Toward Capturing Genetic Epistasis from Multivariate Genome-Wide Association Studies Using Mixed-Precision Kernel Ridge Regression.”
    International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), 2024. Gordon Bell Prize Finalist.
  8. J. Ren, B. Liu, X. Feng, X. Pan, Y. Fu, Y. Yang, and L. Mai. “TorchOpt: An Efficient Library for Differentiable Optimization.”
    Journal of Machine Learning Research (JMLR), 2023.
  9. B. Liu, X. Feng, J. Ren, L. Mai, R. Zhu, H. Zhang, J. Wang, and Y. Yang. “A Theoretical Understanding of Gradient Bias in Meta-Reinforcement Learning.”
    Conference on Neural Information Processing Systems (NeurIPS), 2022.
  10. X. Xia, W. Yang, J. Ren, Y. Li, Y. Zhan, B. Han, and T. Liu. “Pluralistic Image Completion with Probabilistic Mixture-of-Experts.”
    Conference on Neural Information Processing Systems (NeurIPS), 2022.
  11. J. Ren, W. Liang, R. Yan, L. Mai, S. Liu, and X. Liu. “MegBA: A High-Performance and Distributed Library for Large-Scale Bundle Adjustment.”
    European Conference on Computer Vision (ECCV), 2022.
  12. L. Tian, B. Chen, J. Ren, H. Zhang, Z. Wu, N. Han, Y. Chen, and H. Liu. “Multi-Scale Visual Attention for Attribute Disambiguation in Zero-Shot Learning.”
    Signal Processing: Image Communication, 2021.
  13. Z. Duan, D. Wang, B. Chen, C. Wang, W. Chen, Y. Li, J. Ren, and M. Zhou. “Sawtooth Factorial Topic Embeddings Guided Gamma Belief Network.”
    International Conference on Machine Learning (ICML), 2021.

Preprints

  • J. Ren, Y. Li, Z. Ding, W. Pan, and H. Dong. “Probabilistic Mixture-of-Experts for Efficient Deep Reinforcement Learning.”
    arXiv:2104.09122, 2021.

Selected Academic Activities

  • Invited Participant, Dagstuhl Seminar 26392 (September 2026)
  • Young Researcher, 13th Heidelberg Laureate Forum (September 2026)

Service

Conference Reviewer

  • International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), 2026
  • IEEE International Conference on Robotics and Automation (ICRA), 2025
  • International Conference on Parallel Processing (ICPP), 2024
  • International Conference on Learning Representations (ICLR), 2023–2025
  • Conference on Neural Information Processing Systems (NeurIPS), 2022–2023
  • International Conference on Machine Learning (ICML), 2022–2024
  • IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2021–2022

Journal Reviewer

  • ACM Transactions on Mathematical Software (ACM TOMS)
  • Parallel Computing
  • International Journal of Computer Vision (IJCV)
  • IEEE Robotics and Automation Letters (IEEE RA-L)

Awards

  • ACM/IEEE-CS George Michael Memorial High Performance Computing Fellowship (2026)
  • Gordon Bell Prize Finalist, SC (2024)
  • KAUST Dean's Award (2024)
  • 1st Place, ISC Student Cluster Competition Coding Challenge (2024)
  • HiPEAC Travel & Collaboration Grant (2023)