Development of a Virtual AI Testbed Capable of Performance Verification Before Building Massive AI Servers
< From left: Professor Jongse Park , M.S candidate Jaehong Cho, M.S candidate Hyunmin Choi, Professor Brandon Reagen ISPASS >
Operating Large Language Model (LLM) services like ChatGPT requires a server infrastructure on the scale of tens of thousands of units. However, constructing actual equipment every time a new AI semiconductor or system architecture needs to be verified incurs massive costs and time. A research team at our university has developed a ‘virtual testbed’ that can pre-verify performance and efficiency inside a computer before building an actual large-scale AI server.
KAIST announced on May 29th that the research on a Large Language Model (LLM) serving infrastructure simulator (virtual testing software) developed by Professor Jongse Park’s research team in the School of Computing won the Best Paper Award at ‘ISPASS 2026 (IEEE International Symposium on Performance Analysis of Systems and Software),’ a world-renowned conference in the field of computer system performance analysis.
‘LLMServingSim 2.0,’ developed by the research team, is a simulation platform capable of virtually analyzing various hardware and software combinations in complex AI service environments. Researchers and developers can freely experiment with various design options and verify performance without having to directly build expensive, large-scale server infrastructures.
< LLMServingSim 2.0 is workload >
In particular, this technology is drawing attention because it goes beyond the existing Graphics Processing Unit (GPU)-centric environment to support diverse hardware environments, including Neural Processing Units (NPUs), which are rising as next-generation AI semiconductors, and Processing-In-Memory (PIM, a semiconductor technology that performs operations inside the memory).
In other words, it is a technology that allows future-oriented AI semiconductors that have not yet been commercialized to be tested in advance within a virtual datacenter environment. Through this, it is possible to replicate and analyze inside a computer how much the service speed improves, how much power consumption is reduced, and whether it operates stably even in a server environment scaled to tens of thousands of units when a specific semiconductor is applied.
In addition, it reproduces complex operations that occur during actual AI service operations—such as data processing, request distribution, and memory utilization—at the system level, enabling performance evaluations that are close to reality. Notably, it can even analyze disaggregated infrastructure environments where multiple server resources are separated and connected for use, showing great potential for utilization in next-generation AI datacenter research.
This simulator is expected to be widely utilized not only by researchers but also by LLM service companies and AI semiconductor startups to design and optimize next-generation AI infrastructures. This is because it can rapidly verify new AI semiconductors or service architectures prior to actual construction, thereby significantly reducing the cost and time of AI infrastructure development.
< Research Image (AI-generated image) >
Professor Jongse Park said, “The competitiveness of AI services is determined not only by the model itself but also by the infrastructure technology that operates it stably and efficiently.” He added, “We hope this simulator will serve as an important foundation for researchers and the industry to develop next-generation AI infrastructures faster and more efficiently.”
This research was led by M.S candidate Jaehong Cho and Hyunmin Choi in the School of Computing as co-first authors. Following their Best Paper Award at the 2024 IISWC (IEEE International Symposium on Workload Characterization), the research team won the Best Paper Award again at this ISPASS 2026, proving their research competitiveness in the field of AI infrastructure once more.
※ Paper Title: LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure, DOI: 10.1109/ISPASS69572.2026.00012 (Authors: Jaehong Cho, Hyunmin Choi, Guseul Heo, Jongse Park) ※ Open Source Link: https://llmservingsim.ai/ Meanwhile, this research was conducted with support from the Ministry of Science and ICT (MSIT), the Institute for Information & Communications Technology Planning & Evaluation (IITP, No. RS-2024-00396013), the Electronics and Telecommunications Research Institute (ETRI, No. RS-2025-02305453), and SK hynix.
KAIST Team Wins Grand Prize at Kakao AI Incubation Project
<(From Left) Professor Jongse Park, Professor youngjin Kwon, Professor Jaehyuk Huh, Professor Knunle Olukotun>
Currently, Large Language Model (LLM) services like ChatGPT rely heavily on expensive GPU servers. This structure faces significant limitations, as costs and power consumption skyrocket as service scales increase. Researchers at KAIST have developed a next-generation AI infrastructure technology to address these challenges.
KAIST announced on January 30th that the ‘AnyBridge AI’ team, led by Professor Jongse Park from the School of Computing, has developed a next-generation AI infrastructure software. This software allows for efficient LLM services by integrating various AI accelerators instead of relying solely on GPUs. The technology won the Grand Prize at the "4 ISTs (Science & Tech Institutes) × Kakao AI Incubation Project" hosted by Kakao.
This project is a joint industry-academic collaboration between Kakao and the four major science and technology institutes (KAIST, GIST, DGIST, and UNIST). It selected outstanding teams by evaluating the technical prowess and business viability of preliminary startup teams based on AI technology. The Grand Prize winning team receives a total of 20 million KRW in prize money and up to 35 million KRW in Kakao Cloud credits.
AnyBridge AI is a technical startup team led by Professor Jongse Park (CEO), with Professors Youngjin Kwon and Jaehyuk Huh from KAIST's School of Computing participating. Based on research achievements in AI systems and computer architecture, the team aims to develop technology applicable to actual industrial sites. Furthermore, Professor Kunle Olukotun of Stanford University—co-founder of the Silicon Valley AI semiconductor startup SambaNova—is participating as an advisor to push for global technology and business expansion.
The AnyBridge team noted that most current LLM services are dependent on expensive GPU infrastructure, leading to structural limits where operating costs and power usage surge as services scale. The researchers analyzed that the root cause of this issue lies not in the performance of specific hardware, but in the absence of a system software layer capable of efficiently connecting and operating various AI accelerators, such as NPUs (AI-specialized chips) and PIMs (next-gen chips that process AI within memory), alongside GPUs.
<Technical diagram of AnyBridge: Enhancing LLM performance by flexibly utilizing various AI accelerators>
In response, the AnyBridge team proposed an integrated software stack that can service LLMs across the same interface and runtime environment, regardless of the accelerator type. Specifically, they received high praise for pointing out the limitations of existing GPU-centric LLM serving structures and presenting a "Multi-Accelerator LLM Serving Runtime Software" as their core technology.
This technology enables the implementation of a flexible AI infrastructure where the most suitable AI accelerator can be selected and combined based on the task's characteristics, without being tied to a specific vendor or hardware. This is evaluated as a major advantage that can reduce costs and power consumption while significantly increasing scalability for LLM services.
<Illustration of the Multi-Accelerator LLM Service Platform - AI-generated image>
Additionally, based on years of accumulated research in LLM serving system simulation, the AnyBridge team possesses a research foundation that can pre-verify various hardware/software design combinations without building a large-scale physical infrastructure. This point demonstrated both the technical maturity and the industrial feasibility of their work.
"This award is a result of recognizing the necessity of system software that integrates various AI accelerators, moving beyond the limits of GPU-centric AI infrastructure," said Professor Jongse Park. He added, "It is meaningful that we could expand our research results into industrial fields and entrepreneurship. We will continue to develop this into a core technology for next-generation LLM serving infrastructure through cooperation with industrial partners."
This award is seen as a prime example of KAIST's research moving beyond academic papers toward next-generation AI infrastructure technology and startups. AnyBridge AI plans to advance and verify its technology through future collaborations with Kakao and related industrial partners.
<Photo of the Grand Prize ceremony: Left - Kakao Investment CEO Do-young Kim; Right - KAIST Prof. Jongse Park>
KAIST Develops an AI Semiconductor Brain Combining Transformer's Intelligence and Mamba's Efficiency
<(From Left) Ph.D candidate Seongryong Oh, Ph.D candidate Yoonsung Kim, Ph.D candidate Wonung Kim, Ph.D candidate Yubin Lee, M.S candidate Jiyong Jung, Professor Jongse Park, Professor Divya Mahajan, Professor Chang Hyun Park>
As recent Artificial Intelligence (AI) models’ capacity to understand and process long, complex sentences grows, the necessity for new semiconductor technologies that can simultaneously boost computation speed and memory efficiency is increasing. Amidst this, a joint research team featuring KAIST researchers and international collaborators has successfully developed a core AI semiconductor 'brain' technology based on a hybrid Transformer and Mamba structure, which was implemented for the first time in the world in a form capable of direct computation inside the memory, resulting in a four-fold increase in the inference speed of Large Language Models (LLMs) and a 2.2-fold reduction in power consumption.
KAIST (President Kwang Hyung Lee) announced on the 17th of October that the research team led by Professor Jongse Park from KAIST School of Computing, in collaboration with Georgia Institute of Technology in the United States and Uppsala University in Sweden, developed 'PIMBA,' a core technology based on 'AI Memory Semiconductor (PIM, Processing-in-Memory),' which acts as the brain for next-generation AI models.
Currently, LLMs such as ChatGPT, GPT-4, Claude, Gemini, and Llama operate based on the 'Transformer' brain structure, which sees all of the words simultaneously. Consequently, as the AI model grows and the processed sentences become longer, the computational load and memory requirements surge, leading to speed reductions and high energy consumption as major issues.
To overcome these problems with Transformer, the recently proposed sequential memory-based 'Mamba' structure introduced a method for processing information over time, increasing efficiency. However, memory bottlenecks and power consumption limits still remained.
Professor Park Jongse's research team designed 'PIMBA,' a new semiconductor structure that directly performs computations inside the memory in order to maximize the performance of the 'Transformer–Mamba Hybrid Model,' which combines the advantages of both Transformer and Mamba.
While existing GPU-based systems move data out of the memory to perform computations, PIMBA performs calculations directly within the storage device without moving the data. This minimizes data movement time and significantly reduces power consumption.
<Analysis of Post-Transformer Models and Proposal of a Problem-Solving Acceleration System>
As a result, PIMBA showed up to a 4.1-fold improvement in processing performance and an average 2.2-fold decrease in energy consumption compared to existing GPU systems.
The research outcome is scheduled to be presented on October 20th at the '58th International Symposium on Microarchitecture (MICRO 2025),' a globally renowned computer architecture conference that will be held in Seoul. It was previously recognized for its excellence by winning the Gold Prize at the '31st Samsung Humantech Paper Award.' ※Paper Title: Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving, DOI: 10.1145/3725843.3756121
This research was supported by the Institute for Information & Communications Technology Planning & Evaluation (IITP), the AI Semiconductor Graduate School Support Project, and the ICT R&D Program of the Ministry of Science and ICT and the IITP, with assistance from the Electronics and Telecommunications Research Institute (ETRI). The EDA tools were supported by IDEC (the IC Design Education Center).
KAIST Develops Novel Candidiasis Treatment Overcoming Side Effects and Resistance
<(From left) Ph. D Candidate Ju Yeon Chung, Prof.Hyun Jung Chung, Ph.D candidate Seungju Yang, Ph.D candidate Ayoung Park, Dr. Yoon-Kyoung Hong from Asan Medical Center, Prof. Yong Pil Chong, Dr. Eunhee Jeon>
Candida, a type of fungus, which can spread throughout the body via the bloodstream, leading to organ damage and sepsis. Recently, the incidence of candidiasis has surged due to the increase in immunosuppressive therapies, medical implants, and transplantation. Korean researchers have successfully developed a next-generation treatment that, unlike existing antifungals, selectively acts only on Candida, achieving both high therapeutic efficacy and low side effects simultaneously.
KAIST (President Kwang Hyung Lee) announced on the 8th that a research team led by Professor Hyun-Jung Chung of the Department of Biological Sciences, in collaboration with Professor Yong Pil Jeong's team at Asan Medical Center, developed a gene-based nanotherapy (FTNx) that simultaneously inhibits two key enzymes in the Candida cell wall.
Current antifungal drugs for Candida have low target selectivity, which can affect human cells. Furthermore, their therapeutic efficacy is gradually decreasing due to the emergence of new resistant strains. Especially for immunocompromised patients, the infection progresses rapidly and has a poor prognosis, making the development of new treatments to overcome the limitations of existing therapies urgent.
The developed treatment can be administered systemically, and by combining gene suppression technology with nanomaterial technology, it effectively overcomes the structural limitations of existing compound-based drugs and successfully achieves selective treatment against only Candida.
The research team created a gold nanoparticle-based complex loaded with short DNA fragments called antisense oligonucleotides (ASO), which simultaneously target two crucial enzymes—β-1,3-glucan synthase (FKS1) and chitin synthase (CHS3)—important for forming the cell wall of the Candida fungus.
By applying a surface coating technology that binds to a specific glycolipid structure (a structure combining sugar and fat) on the Candida cell wall, a targeted delivery device was implemented. This successfully achieved a precise targeting effect, ensuring the complex is not delivered to human cells at all but acts selectively only on Candida.
<Figure 1: Overview of antifungal therapy design and experimental approach>
This complex, after entering Candida cells, cleaves the mRNA produced by the FKS1 and CHS3 genes, thereby inhibiting translation and simultaneously blocking the synthesis of cell wall components β-1,3-glucan and chitin. As a result, the
Candida cell wall loses its structural stability and collapses, suppressing bacterial survival and proliferation.
In fact, experiments using a systemic candidiasis model in mice confirmed the therapeutic effect: a significant reduction in
Candida count in the organs, normalization of immune responses, and a notable increase in survival rates were observed in the treated group.
Professor Hyun-Jung Chung, who led the research, stated, "This study presents a method to overcome the issues of human toxicity and drug resistance spread with existing treatments, marking an important turning point by demonstrating the applicability of gene therapy for systemic infections". She added, "We plan to continue research on optimizing administration methods and verifying toxicity for future clinical application."
This research involved Ju Yeon Chung and Yoon-Kyoung Hong as co-first authors , and was published in the international journal 'Nature Communications' on July 1st.
Paper Title: Effective treatment of systemic candidiasis by synergistic targeting of cell wall synthesis
DOI: 10.1038/s41467-025-60684-7
This research was supported by the Ministry of Health and Welfare and the National Research Foundation of Korea.
Development of Core NPU Technology to Improve ChatGPT Inference Performance by Over 60%
Latest generative AI models such as OpenAI's ChatGPT-4 and Google's Gemini 2.5 require not only high memory bandwidth but also large memory capacity. This is why generative AI cloud operating companies like Microsoft and Google purchase hundreds of thousands of NVIDIA GPUs. As a solution to address the core challenges of building such high-performance AI infrastructure, Korean researchers have succeeded in developing an NPU (Neural Processing Unit)* core technology that improves the inference performance of generative AI models by an average of over 60% while consuming approximately 44% less power compared to the latest GPUs.
*NPU (Neural Processing Unit): An AI-specific semiconductor chip designed to rapidly process artificial neural networks.
On the 4th, Professor Jongse Park's research team from KAIST School of Computing, in collaboration with HyperAccel Inc. (a startup founded by Professor Joo-Young Kim from the School of Electrical Engineering), announced that they have developed a high-performance, low-power NPU (Neural Processing Unit) core technology specialized for generative AI clouds like ChatGPT.
The technology proposed by the research team has been accepted by the '2025 International Symposium on Computer Architecture (ISCA 2025)', a top-tier international conference in the field of computer architecture.
The key objective of this research is to improve the performance of large-scale generative AI services by lightweighting the inference process, while minimizing accuracy loss and solving memory bottleneck issues. This research is highly recognized for its integrated design of AI semiconductors and AI system software, which are key components of AI infrastructure.
While existing GPU-based AI infrastructure requires multiple GPU devices to meet high bandwidth and capacity demands, this technology enables the configuration of the same level of AI infrastructure using fewer NPU devices through KV cache quantization*. KV cache accounts for most of the memory usage, thereby its quantization significantly reduces the cost of building generative AI clouds.
*KV Cache (Key-Value Cache) Quantization: Refers to reducing the data size in a type of temporary storage space used to improve performance when operating generative AI models (e.g., converting a 16-bit number to a 4-bit number reduces data size by 1/4).
The research team designed it to be integrated with memory interfaces without changing the operational logic of existing NPU architectures. This hardware architecture not only implements the proposed quantization algorithm but also adopts page-level memory management techniques* for efficient utilization of limited memory bandwidth and capacity, and introduces new encoding technique optimized for quantized KV cache.
*Page-level memory management technique: Virtualizes memory addresses, as the CPU does, to allow consistent access within the NPU.
Furthermore, when building an NPU-based AI cloud with superior cost and power efficiency compared to the latest GPUs, the high-performance, low-power nature of NPUs is expected to significantly reduce operating costs.
Professor Jongse Park stated, "This research, through joint work with HyperAccel Inc., found a solution in generative AI inference lightweighting algorithms and succeeded in developing a core NPU technology that can solve the 'memory problem.' Through this technology, we implemented an NPU with over 60% improved performance compared to the latest GPUs by combining quantization techniques that reduce memory requirements while maintaining inference accuracy, and hardware designs optimized for this".
He further emphasized, "This technology has demonstrated the possibility of implementing high-performance, low-power infrastructure specialized for generative AI, and is expected to play a key role not only in AI cloud data centers but also in the AI transformation (AX) environment represented by dynamic, executable AI such as 'Agentic AI'."
This research was presented by Ph.D. student Minsu Kim and Dr. Seongmin Hong from HyperAccel Inc. as co-first authors at the '2025 International Symposium on Computer Architecture (ISCA)' held in Tokyo, Japan, from June 21 to June 25. ISCA, a globally renowned academic conference, received 570 paper submissions this year, with only 127 papers accepted (an acceptance rate of 22.7%).
※Paper Title: Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
※DOI: https://doi.org/10.1145/3695053.3731019
Meanwhile, this research was supported by the National Research Foundation of Korea's Excellent Young Researcher Program, the Institute for Information & Communications Technology Planning & Evaluation (IITP), and the AI Semiconductor Graduate School Support Project.