Managing and directing research projects on audio-visual understanding, generation, and evaluation
- PI
- Prof. Se-Young Yun & Prof. Tae-Hyun Oh, KAIST
Managing and directing research projects on audio-visual understanding, generation, and evaluation
Worked on open-world visual document intelligence distilled from LLMs and document expert models
Worked as a Student Researcher in the semiconductor wafer cleaning process team under Future Technology Research Institute, SK Hynix
Teaching and implementation for computer vision tasks
Review, implementation, and demo for papers on pruning and knowledge distillation
| Title | First Author | Year | Venue |
|---|---|---|---|
| Learning both Weights and Connections for Efficient Neural Networks | Song Han | 2015 | NeurIPS |
| Pruning Filters for Efficient ConvNets | Hao Li | 2017 | ICLR |
| The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks | Jonathan Frankle | 2019 | ICLR |
| SNIP: Single-shot Network Pruning based on Connection Sensitivity | Namhoon Lee | 2019 | ICLR |
| Comparing Rewinding and Fine-tuning in Neural Network Pruning | Alex Renda | 2020 | ICLR |
| The Generalization-Stability Tradeoff In Neural Network Pruning | Brian R. Bartoldson | 2020 | NeurIPS |
| Title | First Author | Year | Venue |
|---|---|---|---|
| Distilling the Knowledge in a Neural Network | Geoffrey Hinton | 2015 | NeurIPS |
| Knowledge Distillation by On-the-Fly Native Ensemble | Xu Lan | 2018 | NeurIPS |
| Relational Knowledge Distillation | Wonpyo Park | 2019 | CVPR |
| Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation | Linfeng Zhang | 2019 | ICCV |
| Regularizing Class-wise Predictions via Self-knowledge Distillation | Sukmin Yun | 2020 | CVPR |
| DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter | Victor Sanh | 2020 | NeurIPS, Workshop |
Leading a research group developing and evaluating audio-driven talking-avatar generation models
Audio-driven talking avatars must be judged not only by how accurately the lips follow the speech, but also by whether the face fully expresses the intended emotion. This project develops audio-driven talking-avatar generation models and studies how to evaluate them reliably and quantitatively. We examine whether multimodal LLMs can serve as judges of the emotion expressed by generated 3D talking heads and investigate how they perceive different visual representations of rendered videos.
Led a research group developing collaborative federated learning algorithms
Federated learning is expected to become a foundational technology for embodied AI and robotics, where multiple agents need to learn collaboratively. In such multi-agent systems, each agent collects its own data in a distributed manner, and the agents must learn together from the data that is reliable. This project is aimed at developing a collaborative learning framework for multi-agent systems, focused on federated learning within the context of data-heterogeneity among multiple agents. We aim to develop both general and personalized federated learning algorithms that surpass existing state-of-the-art solutions.
A significant part of this project is classified as confidential. Only publicly disclosed parts were included in this content.
The global pandemic of COVID-19, first reported in China in December 2019, is causing many people to suffer. This pandemic caused each government to make mistakes of increasing the damage and scale without responding properly. The existing statistical or SIR spread model is still used by government quarantine officials to measure spread rate of infectious diseases. However, there is a lack of consideration for the actual impact of each policy on the spread of infectious diseases between regions or countries. In addition, various factors such as people’s mobility change or distancing policies are difficult to be considered with SIR models. In order to establish and implement prevention policies early, it is necessary to predict the spread of infectious diseases and understand the influence of policies by considering various complex factors. In this study, we develop an early prediction model for the spread of infectious diseases, based on explainable artificial intelligence (XAI) to provide a basis for responding properly in future crisis situations.
Outputs: Technical report, Code
The global pandemic of COVID-19 that has continued from 2019 to the present indicates the importance of early diagnosis of infectious diseases. If we can detect in advance that a disease has a high risk of spreading significantly, it will not only be able to specify information that needs to be searched, but also prevent the spread of infectious diseases by appropriate early responses. In this project, we develop an infectious disease monitoring/surveillance model, so that important information is analyzed to quickly grasp the risk in the early stages of infectious diseases. Also, by integrating the monitoring model and the spreading model, we enable complex analysis of infectious diseases.
Outputs: Code
PMLR Workshop 2022(Oral)
Real-time and Explainable Detection of Epidemics with Global News DataWe are often unavailable to obtain enough labeled data for our target task (e.g., image classification) in most real-world cases. Some recent works have addressed the data shortage by proposing un-/semi-/self-supervised learning methods that utilize unlabeled data to improve the performance in target task without or with only a few amounts of labeled data. In this project, we extend those methods to build a data and training efficient algorithm to cope with various real-world constraints.
Outputs: Code
It is crucial for the model not only to accurately predict, but also to predict how certain it is. If we can find areas where the model's predictions are uncertain, it can greatly reduce the labor force because a person can intervene and correct only those parts. In the semiconductor industry, AI models (semantic segmentation models) segment wafer photos for the quality control of semiconductor wafers, but the amount is too large and ultra-high resolution, making it difficult for humans to inspect them one by one. After the segmentation model makes a prediction, if only the unclear parts are given to a person, he/she can correct those few samples. Therefore, in this work, we intend to create a model that filters out uncertain data by introducing a learning methodology that predicts how certain the prediction of the image segmentation model is.
Outputs: Code
Developed an uncertainty-aware semi-supervised semantic segmentation model for safety-critical applications
Semantic segmentation models used in safety-critical applications such as autonomous driving need to account for how uncertain their predictions are. In this project, we study uncertainty modeling and calibration for semantic segmentation and develop an uncertainty-aware semi-supervised model, using video data recorded from real vehicles.