TY - GEN
T1 - Personalized Learning Path Planning through Goal-Driven Learner State Modeling
AU - Lim, Joy Jia Yin
AU - He, Ye
AU - Yu, Jifan
AU - Cong, Xin
AU - Zhang-Li, Daniel
AU - Liu, Zhiyuan
AU - Liu, Huiqin
AU - Hou, Lei
AU - Li, Juanzi
AU - Xu, Bin
N1 - Publisher Copyright:
© 2026 Owner/Author.
PY - 2026/4/12
Y1 - 2026/4/12
N2 - Personalized Learning Path Planning (PLPP) aims to design adaptive learning paths that align with individual goals. While large language models (LLMs) show potential in personalizing learning experiences, existing approaches often lack mechanisms for goal-aligned planning. We introduce Pxplore, a novel framework for PLPP that integrates a reinforcement-based training paradigm and an LLM-driven educational architecture. We design a structured learner state model and an automated reward function that transforms abstract objectives into computable signals. We train the policy combining supervised fine-tuning (SFT) and Group Relative Policy Optimization (GRPO), and deploy it within a real-world learning platform. Extensive experiments validate Pxplore's effectiveness in producing coherent, personalized, and goal-driven learning paths. We release our code and dataset at https://github.com/Pxplore/pxplore-algo.
AB - Personalized Learning Path Planning (PLPP) aims to design adaptive learning paths that align with individual goals. While large language models (LLMs) show potential in personalizing learning experiences, existing approaches often lack mechanisms for goal-aligned planning. We introduce Pxplore, a novel framework for PLPP that integrates a reinforcement-based training paradigm and an LLM-driven educational architecture. We design a structured learner state model and an automated reward function that transforms abstract objectives into computable signals. We train the policy combining supervised fine-tuning (SFT) and Group Relative Policy Optimization (GRPO), and deploy it within a real-world learning platform. Extensive experiments validate Pxplore's effectiveness in producing coherent, personalized, and goal-driven learning paths. We release our code and dataset at https://github.com/Pxplore/pxplore-algo.
KW - group relative policy optimization
KW - learner state modeling
KW - personalized learning path planning
KW - reinforcement learning
UR - https://www.scopus.com/pages/publications/105038570529
U2 - 10.1145/3774904.3792245
DO - 10.1145/3774904.3792245
M3 - 会议稿件
AN - SCOPUS:105038570529
T3 - WWW 2026 - Proceedings of the ACM Web Conference 2026
SP - 6067
EP - 6078
BT - WWW 2026 - Proceedings of the ACM Web Conference 2026
PB - Association for Computing Machinery, Inc
T2 - 35th ACM Web Conference, WWW 2026
Y2 - 29 June 2026 through 3 July 2026
ER -