Tianrui Song's PhD research will develop responsible multimodal AI tools to support video creators in managing and editing footage. The project will use large language models, vision-language models, and video-language models to build narrative-aware representations of video clips, capturing not only visual content but also editorial function, such as A-roll, B-roll, reactions, transitions, and establishing shots. These representations will support creator-adaptive retrieval and recommendation over personal footage libraries, allowing users to search and refine results through natural language and timeline context. The project will also develop an interactive editing assistant and evaluate it with real users, measuring efficiency, trust, perceived control, and creative support. The research contributes to human-centred, explainable, and responsible AI for creative workflows.