×

News

[ECE Department] Professor Jaeyoung Do’s research team accepted for Oral Presentation at ICML 2026 for research on human value-based LLM alignment

July 1, 2026l Hit 111

Professor Jaeyoung Do’s research team at the AIDAS Lab in the Department of Electrical and Computer Engineering at Seoul National University announced that its paper, “VALUEFLOW,” which studies the alignment of large language models based on human values, has been accepted as an Oral Presentation at the International Conference on Machine Learning (ICML) 2026. ICML is one of the world’s most prestigious conferences in artificial intelligence and machine learning. At ICML 2026, which will be held in Seoul, Oral Presentation is a presentation format granted only to outstanding research, accounting for approximately 0.7% of all submitted papers.

Figure 1. Architecture of VALUEFLOW

 

Large language models are increasingly being used in a wide range of high-risk and high-value domains, including education, healthcare, policy, and decision-making support. As a result, technologies that allow models to understand and adjust the human values and judgment criteria underlying their responses are becoming increasingly important, beyond simply following user preferences.

The Need for Human Value-Based AI Alignment

 

Existing AI alignment research has primarily adjusted models based on user preferences or feedback scores. However, preferences can easily change depending on how a question is asked or the context in which it is presented, making it difficult to stably reflect the values that individuals or groups fundamentally consider important. For example, even for the same issue, some users may prioritize fairness, while others may place greater importance on freedom or safety.

 

To address these limitations, the research team proposed VALUEFLOW, an integrated framework that can represent, measure, and steer the responses of large language models from the perspective of human values.

 

VALUEFLOW: An Integrated Framework Connecting Value Representation, Measurement, and Steering

 

VALUEFLOW consists of three main stages. First, in the Value Representation stage, the framework integrates different value theories into a single structured embedding space through HiVES, a hierarchical value embedding model. This enables the model to capture value signals at the text level, including care, fairness, freedom, safety, rights, and responsibility, as defined in value systems such as Schwartz’s theory of basic values, moral foundations theory, and rights- and duty-based value frameworks.

 

Second, in the Value Measurement stage, VALUEFLOW uses the large-scale Value Intensity Database (VIDB) and a comparison-based evaluation method to quantitatively assess not only whether a particular value is present, but also how strongly it is expressed. Third, in the Value Steering stage, the framework guides model responses to reflect specific value directions and intensities, while analyzing the steerability and limitations of each model.

Figure 2. Comparison of value steerability across major LLMs. The figure shows differences in model responses to positive and negative value steering.

 

Analyzing the Value Steerability of 10 Major Large Language Models

 

Using VALUEFLOW, the research team evaluated the steerability of 10 major large language models, including GPT-4.1, Gemini, Claude, Qwen, Mistral, Gemma, and Grok, across 4 value theories and 32 value categories.

 

The analysis revealed clear differences in value-steering capabilities across models. Some models responded well to positively reinforcing certain values but showed little response to steering in a negative direction. In particular, values generally considered socially desirable, such as care and universalism, showed strong resistance to negative steering.

 

The team also found that when multiple values were steered simultaneously, similar values tended to be reinforced together, while conflicting values tended to weaken each other’s expression. This suggests that value alignment in large language models is not merely a matter of following instructions, but is closely connected to the models’ internal safety and alignment characteristics.

 

VALUEFLOW can be applied to personalized AI, culture-specific AI alignment, policy-sensitive AI deployment, model auditing, and real-time dialogue-based alignment. In particular, because it can analyze which values a model reflects well and which values it resists, VALUEFLOW is regarded as a foundational technology for transparent and responsible AI development.

 

This research was led by Woojin Kim of the Department of Electrical and Computer Engineering at Seoul National University as the first author, with Sieun Hyeon and Jusang Oh as co-authors and Professor Jaeyoung Do as the corresponding author. The research team plans to expand VALUEFLOW to multi-turn dialogue-based personalized alignment, culture-specific value profiling, and multimodal AI alignment.

 

Professor Do stated, “This research is meaningful because it presents a direction for large language models to move beyond simply following users’ immediate preferences and toward structurally understanding and steering the values that humans consider important. We will continue to develop this work into responsible AI alignment technology that enables the coexistence of diverse individual and societal values.”

 

Source: https://ece.snu.ac.kr/ece/news?md=v&bbsidx=57850

Translated by: Changhoon Kang, English Editor of the Department of Electrical and Computer Engineering, changhoon27@snu.ac.kr