Erzhen Hu, Yanhe Chen, Mingyi Li, Vrushank Phadnis, Pingmei Xu, Xun Qian, Alex Olwal, David Kim, Seongkook Heo, Ruofei Du
DialogLab: Authoring, Simulating, and Testing Dynamic Human-AI Group Conversations
(Abstract) Designing compelling multi-party conversations involving both humans and AI agents presents significant challenges, particularly in balancing scripted structure with emergent, human-like interactions. We introduce DialogLab, a prototyping toolkit for authoring, simulating, and testing hybrid human-AI dialogues. DialogLab provides a unified interface to configure conversational scenes, define agent personas, manage group structures, specify turn-taking rules, and orchestrate transitions between scripted narratives and improvisation. Crucially, DialogLab allows designers to introduce controlled deviations from the scriptโthrough configurable agents that emulate human unpredictabilityโto systematically probe how conversations adapt and recover. DialogLab facilitates rapid iteration and evaluation of complex, dynamic multi-party human-AI dialogues. An evaluation with both end users and domain experts demonstrates that DialogLab supports efficient iteration and structured verification, with applications in training, rehearsal, and research on social dynamics. Our findings show the value of integrating real-time, human-in-the-loop improvisation with structured scripting to support more realistic and adaptable multi-party conversation design.
(Introduction) The simulation and creation of conversations with embodied agents have long been integral to virtual entertainment environments, especially for non-player characters (NPCs). Beyond gaming, these systems are increasingly deployed in diverse contexts such as education [60], healthcare [77], professional training [47], and retail [53]. Initially constrained by rule-based or scripted mechanisms and limitations in natural language understanding (NLU), the design of embodied agents have more recently leveraged Large Language Models (LLMs) to deliver more adaptive, context-aware dialogues.
Commercial tools like Character.AI1 and Replika2 illustrate this trend, enabling users to create custom AI personalities for one-onone conversations. However, as human-AI interaction continues to evolve, user needs extend beyond single-agent dialogues. In many real-world scenarios, users participate in multi-party conversations, where multiple agents and humans interact under a blend of prescripted and emergent conditions. These settings introduce new complexitiesโusers must navigate shifting roles, manage turn-taking protocols, and negotiate control of the conversation as they move between reactive and proactive stances. For instance, a human trainee may lead with structured questions in a medical simulation [75] but revert to a reactive role when questioned by an AI panelist. In multi-agent contexts, these transitions require sophisticated system support to maintain coherence across overlapping speaking turns, interruptions, and dynamic topic shifts.
Through a formative study, we sought to identify key challenges faced by conversation designers and developers: (1) Simulating how scripted agents balance rigidity and flexibility. (2) Testing interactions with human participants who deviate from expected behaviors. These challenges reveal two core design tensions: First, the authoring and simulating of multi-party hybrid human-AI dialogue systems remain complex, often demanding significant programming or prompt engineering [103], hindering rapid iteration by designers. Even with scripted guidance, balancing narrative structure reflecting diverse group dynamics with unpredictable human input requires extensive testing. Second, designers must navigate a trade-off between structured scripts, which provide consistency and safety, and improvised human-led dialogue, which fosters realism, user agency, and spontaneity [65, 104]. Current approaches often represent two extremes: fully scripted interactions, which ensure control but lack adaptability and spontaneity [55, 108]; and fully generative and automated systems, which allow open-ended dialogue but are difficult to steer toward specific goals [59, 106]. While structure is critical in domains like healthcare or education, rigid scripting can make users feel passive and limit adaptability. Generative systems, by contrast, allow open-ended interaction but lack control for task-oriented or training scenarios. These challenges become even more pronounced in multi-party and group settings due to the increased complexity of managing roles, turntaking, and recovery from interruptions [38, 39, 93].
To address these challenges, we present DialogLab, a unified authoring tool for creating multi-party conversations with hybrid scripted/LLM-driven embodied agents. Our system enables designers to: (1) configure basic persona, scene, and conversation setups that help define diverse group dynamics (Fig. 1a-b); (2) embed LLMdriven agents during prototyping to simulate human unpredictability within pre-defined narratives (Fig. 1c); (3) incorporates features designed to facilitate verification and reflection of these complex group and conversational dynamics.
To evaluate DialogLab, we conducted a user study with five regular users and nine domain experts who used DialogLab to author and test multi-party conversation scenarios. Participants appreciated the diverse conversation flexibilities and agency during the authoring, and reported clear verification of group behaviors, and strong support for real-time improvisation and testing. These findings demonstrate DialogLab โs ability to support structured authoring, dynamic simulation, and reflective verification of complex hybrid human-AI conversations.
In summary, our contributions are:
A flexible framework and authoring paradigm for multi-party human-AI conversations, enabling configuration of group dynamics, social protocols, and hybrid scripted/improvised interactions.
DialogLab, an open-source system3 implementing this framework, supporting the workflow of authoring, simulating, and interactively testing multi-party dialogues.
A human evaluation comparing human agents with reactive agents in group settings and a user study of DialogLab with regular users and domain experts, demonstrating its effectiveness in building and simulating human-AI group conversations.ย
(Conclusion) We present DialogLab, a prototyping toolkit for creation and iterative design of multi-party conversations between humans and AI agents. DialogLab allows users to configure various conversation attributes, including agent roles, interaction patterns, and turn-taking dynamics, to simulate realistic and dynamic dialogues. The toolkit supports both scripted and adaptive conversational elements, offering a hybrid approach that combines the structure of predetermined dialogues with the flexibility of spontaneous interactions. DialogLab aims to streamline the development process by providing an intuitive interface for conversation setup and realtime testing and validation, addressing key challenges such as the complexity of configuration, uncertainty of AI responses, and lack of systematic refinement methods. Our evaluation with 14 participants demonstrates its effectiveness in enhancing user control, reducing development time, and facilitating realistic multi-party conversation simulations across various domains. We hope that DialogLabโs approach and contributions will help advance and inspire the emerging field of Human-AI group conversation dynamics.
Dynamic Human-AI Group Conversations
๋ํ ์ธ์ด ๋ชจ๋ธ(LLM)์ ๋ฐ์ ์ผ๋ก ๊ฐ์ ๋น์, ๋์งํธ ํด๋จผ ๋ฑ ์ธ๊ณต์ง๋ฅ ์์ด์ ํธ์ ์ธ๊ฐ์ด ๊ทธ๋ฃน์ ์ด๋ฃจ์ด ๋ค์๊ฐ ๋ํ(group conversation)๋ฅผ ๋๋๋ ๊ฐ์ ์๋๋ฆฌ์ค(ํฌ์คํฐ ์ธ์ , ๋ชจ์ ๋ฉด์ , ๋์์ธ ๋ฆฌ๋ทฐ ๋ฑ)๊ฐ ํ์ฑํ๋๊ณ ์๋ค.
ํ์ง๋ง ๋ค์ธ์์ด ์ฐธ์ฌํ๋ ๋์ ๋ํ ํ์ดํ๋ผ์ธ์ ๋ฐํ ์์(turn-taking) ์ ์ด, ์ญํ ๋ณํ, ๋ผ์ด๋ค๊ธฐ(interruption) ๋ฑ์ ๋ณต์กํ ์ญํ ๊ด๊ณ๊ฐ ๋์์์ด ๋ฐ์ํ์ฌ ์ด๋ฅผ ์ง๊ด์ ์ผ๋ก ์ค๊ณํ๊ณ ํ
์คํธํ๊ธฐ ์ด๋ ต๋ค.
๊ธฐ์กด ๋ฐฉ์๋ค์ trade-offs
100% ์คํฌ๋ฆฝํธ ๋ฐฉ์: ๋ํ์ ํ๋ฆ๊ณผ ์์๋ฅผ ์๋ฒฝํ๊ฒ ํต์ ํ ์ ์๊ณ ์ผ๊ด์ฑ์ด ๋์ง๋ง, ๋ํ๊ฐ ๋๋ฌด ์ ํํ๋์ด ์ค์ ์ธ๊ฐ ์ฐธ์ฌ์์ ๋๋ฐ ํ๋์ด๋ ์ฆํฅ์ ์ธ ์ง๋ฌธ ๋ฐ ๊ฐ์ ์ ์ ์ฐํ๊ฒ ๋์ฒํ์ง ๋ชปํด ๋ฑ๋ฑํ๊ณ ํ์ค์ฑ์ด ๋จ์ด์ง๋ค.
100% ์์ฑํ(LLM) ๋ฐฉ์: ํ๋กฌํํธ๋ง์ผ๋ก ๋ค์ฑ๋กญ๊ณ ์์ฐ์ค๋ฌ์ด ํ๋ฆฌํ ํน์ ๊ตฌ์ฌํ ์ ์์ง๋ง, ๋ฐํ ์์๊ฐ ๊ผฌ์ด๊ฑฐ๋ AI๊ฐ ์ผ์ฒํฌ๋ก ๋น ์ง๊ธฐ ์ฝ๊ณ (hallucination), ๋์์ด๋๊ฐ ์ํ๋ ํน์ ๋ชฉ์ ์ด๋ ๊ต์ก ๋ชฉํ(scenario goal)๋ก ๋ํ๋ฅผ ์ ๋ํ๊ธฐ ์ด๋ ต๋ค.
๋์์ด๋์ ๋์ ์ง์
์ฅ๋ฒฝ: ๊ธฐ์กด ์ฐ๊ตฌ๋ค์ ์ผ๋์ผ ๋ํ๋ ๊ณ ์ ๋ ์๋๋ฆฌ์ค ๋ด ๊ฐ๋ฐ์ ์น์ค๋์ด ์์๋ค. ๋ค์๊ฐ ๋ํ ์๋๋ฆฌ์ค๋ฅผ ์ค๊ณํ๊ณ ํ
์คํธํ๋ ค๋ฉด ๋ณต์กํ ์ฝ๋ฉ์ด๋ ๋ฐฉ๋ํ ํ๋กฌํํธ ์์ง๋์ด๋ง์ด ํ์ํ์ฌ, ์ ์ ๋ํ ์ฝํ
์ธ ๋ฅผ ์ง์ผ ํ๋ ๊ธฐํ์๋ ๋์์ด๋๋ค์ด ๋น ๋ฅด๊ฒ ๋์์ ์คํ(iteration)ํ๊ธฐ ๊ทน๋๋ก ์ด๋ ค์ด ํ๊ฒฝ์ด๋ค.
DialogLab
์ธ๊ฐ๊ณผ ๋ณต์์ AI ์์ด์ ํธ๋ค์ด ํจ๊ป ์ฐธ์ฌํ๋ ๋ณต์กํ ๋ค์๊ฐ ๋ํ๋ฅผ ์ง๊ด์ ์ผ๋ก ์ค๊ณ(authoring)ํ๊ณ , ์๋ฎฌ๋ ์ด์ (simulating) ๋ฐ ๊ฒ์ฆ(testing)ํ ์ ์๋ ํ๋กํ ํ์ ํดํท์ผ๋ก, ์คํฌ๋ฆฝํธ์ ๊ตฌ์กฐ์ ์์ ์ฑ๊ณผ LLM์ ์ฆํฅ์ฑ(improvisation)์ ๊ฒฐํฉํ๋ค.
ํตํฉ ์ธํฐํ์ด์ค ์ ๊ณต(macro-to-micro): ์ฝ๋ฉ ์์ด๋ ๊ฐ์ ๋ํ ๊ณต๊ฐ(scene)์ ์ค์ ํ๊ณ , ๋ง์ฐ์ค ํด๋ฆญ๊ณผ ์ธ์คํํฐ ์ค์ ์ ํตํด AI ์์ด์ ํธ์ ํ๋ฅด์๋ ์ ์, ๊ทธ๋ฃน ๊ตฌ์กฐ(parties), ๋ฐํ ๊ท์น(turn-taking rules)์ ์ง๊ด์ ์ผ๋ก ๊ตฌ์ฑํ ์ ์๋ค.
ํ์ด๋ธ๋ฆฌ๋ ๋ํ ์กฐ์จ(recovery ์ง์): ์ ํด์ง ์๋๋ฆฌ์ค ํ๋ก์ฐ๋ฅผ ๋ฐ๋ฅด๋ค๊ฐ๋, ์ธ๊ฐ ์ฐธ์ฌ์๊ฐ ๋๋ฐ ๋ฐํ(deviation)๋ฅผ ํ์ ๋ AI ์์ด์ ํธ๋ค์ด ์ฆํฅ์ ์ผ๋ก ๋์ํ๊ณ ๋ค์ ์๋ ์๋๋ฆฌ์ค ๊ถค๋๋ก ์์ฐ์ค๋ฝ๊ฒ ๋ณต๊ท(recovery)ํ ์ ์๋๋ก ์์คํ ์ ์ผ๋ก ์ง์ํ๋ค.
์๋ฎฌ๋ ์ด์
๋ฐ ํ
์คํธ ์์ง(simulated-human): ์ค์ ์ธ๊ฐ์ ๋ฐ๋ ค์ค๊ธฐ ์ , ์ธ๊ฐ์ ์์ธก ๋ถ๊ฐ๋ฅ์ฑ์ ํ๋ด ๋ด๋๋ก ์ค์ ๋ ๊ฒ์ฆ์ฉ AI ์์ด์ ํธ(AI human agent)๋ฅผ ํฌ์
ํ๋ค. ์ด ๊ฐ์ ์ธ๊ฐ์๊ฒ ์ฃผ์ ์ดํ(drift), ์ง๋ฌธ(question), ๊ฐ์ ํํ(emotional) ๋ฑ์ ์ธํ
ํธ๋ฅผ ์ฃผ์ด ์์คํ
์ด ์ผ๋ง๋ ์ ๋ฒํฐ๊ณ ๋ณต๊ตฌ๋๋์ง ์ค์๊ฐ ๋ชจ๋ํฐ๋งํ๊ณ ๊ฒ์ฆ(verification)ํ ์ ์๋ค.
ํ๊ฐ ๋ฐ ์ฃผ์ ๊ธฐ์ฌ์
์ฌ์ฉ์ฑ ๋ฐ ์ ๋ฌธ๊ฐ ํ๋ ํ ์คํธ (N=14): 5๋ช ์ ์ผ๋ฐ ์ฌ์ฉ์์ 9๋ช ์ ๋ํ ๋์์ธ ๋๋ฉ์ธ ์ ๋ฌธ๊ฐ๋ฅผ ๋์์ผ๋ก ์ ์ ์คํฐ๋๋ฅผ ์งํํ ๊ฒฐ๊ณผ, ๋ค์๊ฐ ๋ํ ์์คํ ๊ฐ๋ฐ ์๊ฐ์ ํ๊ธฐ์ ์ผ๋ก ์ค์ด๊ณ ๋ณต์กํ ๊ทธ๋ฃน ํ๋ ๋ฐฉ์๊ณผ ๋ํ ์ญํ์ ๊ตฌ์กฐ์ ์ผ๋ก ๊ฒ์ฆํ ์ ์์์ ์ ์ฆํ๋ค.
์ธ๊ฐ ํต์ ๋ชจ๋(human-control)์ ์ฐ์์ฑ: ์คํ ๊ฒฐ๊ณผ, ์ฌ์ฉ์๊ฐ ์ง์ suggestions ์ค ์ ํํ์ฌ ๊ฐ์ ํ๋ ์ธ๊ฐ ํต์ ๋ชจ๋๊ฐ ๋จ์ ๋ฐ์ํ(reactive)์ด๋ ์์ ์์จํ(autonomous) ๋ชจ๋์ ๋นํด ํต๊ณ์ ์ผ๋ก ์ ์๋ฏธํ๊ฒ ๋์ ๋ชฐ์ ๋(engagement, p < .05)๋ฅผ ๋ณด์ฌ์ฃผ๋ฉฐ ๊ฐ์ฅ ํ์ค์ ์ธ ์๋ฎฌ๋ ์ด์ ๋ชจ๋์์ด ์ฆ๋ช ๋์๋ค.
์๊ท๋ชจ ์๋ฎฌ๋ ์ด์ ์ฐ๊ตฌ (50๊ฐ ๋ํ ๊ฒ์ฆ): 5๊ฐ ์๋๋ฆฌ์ค(WoW ๊ฒ์, ๋์์ธ ๋ฆฌ๋ทฐ, ์์ด์ค๋ธ๋ ์ดํน ๋ฑ)์์ ์ด 50๊ฐ์ ๋ํ ์ธ์คํด์ค๋ฅผ ๋ธ๋ผ์ธ๋ ํ ์คํธํ ๊ฒฐ๊ณผ, ํ๊ฐ์์ 62.5%๊ฐ ์์จ ๊ฐ์ ํ ๋ชจ๋๋ฅผ ์ ํธํ์ผ๋ฉฐ ์ค์ ์ธ๊ฐ ๋ํ ํน์ ์ ๋ฏธ๋ฌํจ(subtle), ์ ๋จธ, ์์ฐ์ค๋ฌ์ด ๋ฌด์ง์ํจ(messiness)์ด ํ๋ฅญํ๊ฒ ํํ๋์๋ค๊ณ ์๋ตํ๋ค.
์คํ์์ค ๊ธฐ์ฌ ๋ฐ ํ์ฉ ๋ถ์ผ: ๋ค์๊ฐ ๋ํ์ ์ฌํ์ ํ๋กํ ์ฝ์ ์ค์ ํ ์ ์๋ ์ํคํ
์ฒ ํ๋ ์์ํฌ๋ฅผ ์ ๋ฆฝํ๊ณ DialogLab ์์คํ
์ ์คํ์์ค๋ก ๊ณต๊ฐํ์ผ๋ฉฐ, ์๋ฃ/์ง๋ฌด ๊ต์ก์ฉ ๋กคํ๋ ์ ์๋ฎฌ๋ ์ด์
, ๊ฒ์ ๋ด NPC ์ญํ ๊ด๊ณ ํ
์คํธ, ๊ฐ์ ํฌ์ปค์ค ๊ทธ๋ฃน ์ธํฐ๋ทฐ ๋ฑ ๋ค์ํ ๋๋ฉ์ธ์ ์ ์ฉ ๊ฐ๋ฅํ๋ค.
์ฅ์
์ฒด๊ณ์ ์ธ ๋ค์๊ฐ ํต์ ๋ ฅ: ๋ง์ดํฌ๋ก ํ๋ฅด์๋ ์ค์ ๊ณผ ๋งคํฌ๋ก ์๋๋ฆฌ์ค ๋ ธ๋๋ฅผ ๊ฒฐํฉํ์ฌ, ๋ํ๊ฐ ๋ชฉํ๋ฅผ ๋ฒ์ด๋์ง ์์ผ๋ฉด์๋ ์๋๊ฐ ๋์น๋ ํ์ด๋ธ๋ฆฌ๋ ๋ํ ํ๊ฒฝ์ ์์ ํ๊ฒ ๊ตฌ์ถํ๋ค.
์์จ์ ํดํ ์ดํน ์กฐ์จ: ์ธ๊ฐ์ด ์ธ์ ๊ฐ์ ํ๋๋ผ๋ ์ค์ผ์คํธ๋ ์ดํฐ๊ฐ ๋งฅ๋ฝ์ ์ค์๊ฐ ๊ณ์ฐํ์ฌ ์ ์ ํ AI ์์ด์ ํธ์๊ฒ ๋ฐํ๊ถ์ ๋๊ธฐ๋ฏ๋ก ํ๋ฆ์ด ๋ถ๋๋ฝ๊ฒ ์ ์ง๋๋ค.
์ฌ์ ์๋ฌ ๋๋ฒ๊น
์ต์ ํ: ๋ํ ๋ถ์ ๋์๋ณด๋๋ฅผ ํตํด ํดํ
์ดํน ๊ท ํ, ๊ฐ์ ๋ณํ, ์ฃผ์ ์ผ๊ด์ฑ์ ์๊ฐํํ์ฌ ์ค์ ์ ์ ํ
์คํธ ์ AI์ ๊ฒฐํจ ๊ตฌ๊ฐ์ ์ ์ ์ ์ผ๋ก ์ก์๋ธ๋ค.
ํ๊ณ์
์ฐ์ฐ ๋น์ฉ ๋ฐ ๋ฏธ์ธ ๋ ์ดํด์: ์ฌ๋ฌ AI ์์ด์ ํธ๊ฐ ๋์์ ๋ํ ์ปจํ ์คํธ์ ํ๋ฅด์๋ ๊ด๊ณ์์ LLM์ ํตํด ์ฐ์ฐํด์ผ ํ๋ฏ๋ก, API ๋น์ฉ์ด ๋ง์ด ๋ค๊ณ ์ฆ๊ฐ์ ์ธ ์ธํฐ๋์ ์ ์ฝ๊ฐ์ ๋ต๋ณ ์ง์ฐ(latency)์ด ๋ฐ์ํ ์ ์๋ค.
ํ๋ฅด์๋ ๊ท ์ผํ(hallucination ์ํ): ๋ํ๊ฐ ์์ญ ํด ์ด์ ๊ธธ์ด์ง๋ฉด ์์ด์ ํธ๋ค์ด ์ด์ ๋ํ ํ์คํ ๋ฆฌ๋ฅผ ๋๋ ๊ณต์ ํ๋ฉด์ ๊ฐ์๊ฐ ๊ฐ์ง ๊ณ ์ ์ ์บ๋ฆญํฐ ๊ฐ์ฑ(sharpness)์ด ํ๋ ค์ง๊ฑฐ๋ ์ด์กฐ๊ฐ ๋์งํ๋๋ ๊ฒฝํฅ์ด ์๋ค.
ํฅํ๊ณผ์
๋ฉํฐ๋ชจ๋ฌ ๋ฐ ํ๋ผ์ธ์ด ํ๋ฌ๊ทธ์ธ: ํ ์คํธ ์ค์ฌ์ ์๋ฎฌ๋ ์ด์ ์ ๋์ด, ์๋ฐํ์ ์์ ์ฒ๋ฆฌ(eye contact), ๋ฏธ์ธ ํ์ (micro-expressions), ๋ชฉ์๋ฆฌ ํค ๋ณ์กฐ ๋ฐ ํ์ด๋ฐ ์ ์ด๋ฅผ ๊ฒฐํฉํ ์์ ํ ๋ฉํฐ๋ชจ๋ฌ ์ ์ ๋๊ตฌ๋ก ํ์ฅํ ๊ฒ์ด๋ค.
์ด๋ํ ๊ทธ๋ฃน ๋ํ ์ค์ผ์ผ๋ง: ์๊ท๋ชจ ๊ทธ๋ฃน์ ๋์ด 100์ธ ์ด์์ ๋๊ท๋ชจ ๊ฐ์ ๋ฏธํ ์ด๋ ๊ฐ์์ค ํ๊ฒฝ์์๋ ๋ ์ดํด์์ ์ค์ผ์คํธ๋ ์ด์ ๋ถ๊ดด ์์ด ๊ฒฌ๊ณ ํ๊ฒ ์๋ํ๋๋ก ์ปจํ ์คํธ ๊ด๋ฆฌ ์์คํ ์ ๊ณ ๋ํํ ์์ ์ด๋ค.
Erzhen Hu, Mingyi Li, Jungtaek Hong, Xun Qian, Alex Olwal, David Kim, Seongkook Heo, Ruofei Du
Thing2Reality: Enabling Spontaneous Creation of 3D Objects from 2D Content using Generative AI in XR Meetings
(Abstract) During remote communication, participants often share both digital and physical content, such as product designs, digital assets, and environments, to enhance mutual understanding. Recent advances in augmented communication have facilitated users to swiftly create and share digital 2D copies of physical objects from video feeds into a shared space. However, conventional 2D representations of digital objects limits spatial referencing in immersive environments. To address this, we propose Thing2Reality, an Extended Reality (XR) meeting platform that facilitates spontaneous discussions of both digital and physical items during remote sessions. With Thing2Reality, users can quickly materialize ideas or objects in immersive environments and share them as conditioned multiview renderings or 3D Gaussians. Thing2Reality enables users to interact with remote objects or discuss concepts in a collaborative manner. Our user studies revealed that the ability to interact with and manipulate 3D representations of objects significantly enhances the efficiency of discussions, with the potential to augment discussion of 2D artifacts.
(Introduction) Shared artifacts, including digital resources (e.g., text, images, videos), and physical objects (e.g., prototypes, printouts), play a crucial role in facilitating effective communication of spatial concepts and the generation of design ideas, especially in creative fields such as product design, architecture, and marketing strategy development. They provide common spatial reference points that bridge gaps between collaborators, enhancing creative exploration and ideation [4]. Besides physical artifacts, designers frequently use online platforms like Google and Pinterest to source relevant digital artifacts that can support their design processes [24]. However, using shared artifacts in remote meetings often pose challenges, especially in scenarios that require quick and spontaneous sharing, such as brainstorming and early stage design sessions. First, artifacts shared via remote meetings are typically in 2D, whether they are captured via camera or retrieved from online repositories, limiting the understanding compared to interactions with physical objects or 3D models. Second, in physical meetings, participants can easily observe and interact with tangible artifacts, which facilitates creative exploration and idea generation processes [4]. However, in remote meetings, this level of interaction is often unavailable or limited.
Several methods have attempted to address these challenges, such as preparing 3D models in advance via CAD or 3D scanning [34], or employing specialized real-time 3D capture setups [62, 78]. While effective, these approaches have limitations: pre-made 3D assets do not support spontaneous sharing, and specialized setups are often impractical for general use. Recent advances in AI-driven text-to-3D and image-to-3D technologies [75] present an accessible and intuitive alternative, lowering barriers to 3D content creation and enabling broader participation in collaborative efforts.
In this paper, we aim to investigate how on-the-fly transformations between 2D content and 3D representations can support object-centric ideation and spatial sense-making in collaborative XR meetings, and how the design of such a system can enhance user experiences by integrating interactive image-to-3D workflows across various XR meeting context.
We designed and implemented Thing2Reality, an XR meeting platform that enables fluid interactions with 2D and 3D artifacts. Thing2Reality allows users to segment content from any source (video streams, shared digital screens) within the XR environment (Figure 1a), generate multi-view renderings (Figure 1b) with conditioned multi-view diffusion models, and transform them into shared 3D Objects with Gaussian splatting for interactive manipulation (Figure 1c). We conducted two user studies: a preliminary study (N=12) analyzing the usability of using digital and physical sources for 3D object creation, and an exploratory study (N=18) exploring how the co-existence and transformation between 2D and 3D formats shape collaboration patterns across tasks such as avatar decoration, spatial layout, and open-ended design brainstorming. Our findings suggest that 3D objects facilitate intuitive explanations, detailed visualization, and interactive collaboration, whereas 2D representations are more often used for final pitch deliverables, suggesting a context-dependent trade-off between the two formats based on task objectives.
In summary, we contribute: Thing2Reality, an XR meeting platform that provides onthe-fly 3D objects generation by enabling users to present and share spontaneous thoughts, and augment their shared digital and physical artifacts with remote collaborators. Findings from a usability study (๐ =12) examining the digital and physical inputs of Thing2Reality. Findings from an exploratory user study (๐ =18) evaluating the use of Thing2Reality (both 2D-to-3D and 3D-to-2D workflow) for discussing and presenting both 2D and 3D objects in XR meetings.
(Conclusion) We believe that XR communication has tremendous promise for co-presence and for bridging distances between humans, yet much focus today is on realistic rendering of avatars and remote participants. However, as XR systems mature and become increasingly realistic, it will also become increasingly important to support a similar level of spontaneity with objects and artifacts, as what people experience in real environments. In this paper, we presented Thing2Reality, an XR communication system that allows users to instantly materialize ideas or physical objects and share them as interactive conditioned multiview renderings or 3D Gaussians for realistic 3D rendering. Thing2Reality is one of many necessary building blocks towards increasingly realistic co-presence in XR, and we hope that our work will inspire continued work towards augmented communication in both physical and mirrored world [9].
XR ์๊ฒฉ ํ์์์ 2D ๊ณต์ ๋ฐฉ์์ trade-offs
2D ๊ณต์ ์ ๋ช ํํ ์ฅ์ : ์น ๋ธ๋ผ์ฐ์ ๊ฒ์ ์ด๋ฏธ์ง, ๋์งํธ ๋ฌธ์, ๊ทธ๋ฆฌ๊ณ ๋ฌผ๋ฆฌ์ ์ธ ์นด๋ฉ๋ผ ํผ๋ ๋ฑ์ ํ์ฉํ์ฌ ํ์ ์ค์ ์์ด๋์ด๋ฅผ ๊ฐ์ฅ ๋น ๋ฅด๊ณ ์ง๊ด์ ์ผ๋ก ์๊ฒฉ ๊ณต๊ฐ์ ๊ฐ์ ธ์ ๋์ธ ์ ์๋ค.
๊ณต๊ฐ ์ ์ด๋ ฅ ๋ถ์กฑ์ ํ๊ณ: ๋ฉํ๋ฒ์ค๋ XR ๊ฐ์ ๋ชฐ์
ํ 3์ฐจ์ ํ๊ฒฝ์์๋ ๋ง์ ๊ณต์ ๋ ์ฝํ
์ธ ๋ ์์ 2D ํ๋ฉด ์นด๋์ ๋ถ๊ณผํฉ๋๋ค. ์ด๋ก ์ธํด ์ ํ ๋์์ธ, ๊ฐ๊ตฌ ๋ฐฐ์น, ๊ฑด์ถ ๋ธ๋ ์ธ์คํ ๋ฐ์ฒ๋ผ 3์ฐจ์ ๊ณต๊ฐ์ ์ฐธ์กฐ(spatial referencing)๋ ๊ฐ๋๋ณ ์ ๋ฐ ์๊ฐํ๊ฐ ํ์์ ์ธ ์์
์์๋, ์์ฌ์ํต์ ์ฌ๊ฐํ ๋ณ๋ชฉ ํ์์ด ๋ฐ์ํ๋ค.
๊ธฐ์กด 3D ๋ชจ๋ธ๋ง ์์ฑ ๋ฐฉ์์ ํ๊ณ
์ฌ์ ์ ์ ๋ฐฉ์(CAD ๋ฐ 3D ์ค์บ): ์ ๋ฐํ 3D assets์ ์ป์ ์ ์์ผ๋, ํ์ ์ ์ ๋ฏธ๋ฆฌ ์ค๋นํด์ผ ํ๋ฏ๋ก ํ์ ๋์ค ๊ฐ์๊ธฐ ๋ ์ค๋ฅธ ์์ด๋์ด๋ฅผ ์ฆ์์์(spontaneous) ๊ณต์ ํ๊ณ ๊ฒ์ฆํ๋ ์๊ฒฉ ์ํฌํ๋ก์ฐ๋ฅผ ์ง์ํ์ง ๋ชปํ๋ค.
์ค์๊ฐ 3D ์บก์ฒ ์ฅ๋น: ๋ค๊ฐ๋ ์นด๋ฉ๋ผ ๋ฆฌ๊ทธ ์ธํธ๋ ๊ณ ๊ฐ์ ํน์ RGB-D ์ผ์ ์ฅ๋น๊ฐ ํ์์ ์ด์ด์, ์ผ๋ฐ์ ์ธ ์ฌ์ฉ์๋ค์ ์ผ์์ ์ธ ์๊ฒฉ ํ์ ํ๊ฒฝ์ ๋์
ํ๊ธฐ์๋ ์ง์
์ฅ๋ฒฝ์ด ๋๋ฌด ๋๊ณ ์ค์ฉ์ฑ์ด ๋จ์ด์ง๋ค.
Thing2Reality
๋น๋์ค ์คํธ๋ฆผ์ด๋ ์น ํ๋ฉด ๊ฐ์ ๋จ ํ ์ฅ์ 2D ์์ค ์ด๋ฏธ์ง์์ ์ฌ์ฉ์๊ฐ ์ํ๋ ๋ฌผ์ฒด๋ง ํก ์๋ผ๋ด์ด(segment), AI ๊ธฐ๋ฐ์ผ๋ก ๋ค๊ฐ๋ ๋ทฐ๋ฅผ ์์ฑํ๊ณ ์ด๋ฅผ ์ฆ์์์ ์กฐ์ํ ์ ์๋ 3D ๊ฐ์ ๋ฌผ์ฒด(3D Gaussian)๋ก ๊ตฌํํด ๋ด๋ ์ค์๊ฐ XR ํ์ ํ๋ซํผ์ด๋ค.
์ค์๊ฐ 2D-to-3D ์ํฌํ๋ก์ฐ: ์น์ํ ์ค์ธ ์ด๋ฏธ์ง๋ ZED Mini ์นด๋ฉ๋ผ ํผ๋๋ก ๋น์ถ ์ค์ ์ฌ๋ฌผ ์์์ VR ์ปจํธ๋กค๋ฌ๋ฅผ ํ์ฉํด ์ฐ์์ ์ธ ์คํธ๋กํฌ(marking)๋ฅผ ๊ทธ๋ฆฌ๋ฉด, MobileSAM ๊ธฐ๋ฐ์ผ๋ก ๊ฐ์ฒด๋ฅผ ์ค์๊ฐ ๋ถ๋ฆฌ(segmentation)ํ์ฌ 3D ๋ชจ๋ธ๋ก ๋๋ฑ ๋ณํํ๋ค.
3D Gaussian Splatting ๊ธฐ์ ์ตํฉ: ํ ์คํธ ์ ๋ ฅํ ์์ฑ ๊ธฐ์ ๊ณผ ๋ฌ๋ฆฌ, ๋ํ ์ฌ๊ตฌ์ฑ ๋ชจ๋ธ(LGM) ๋ฐ ๋ค๊ฐ๋ ํ์ฐ ๋ชจ๋ธ์ ํตํด ์๋ณธ 2D ์ด๋ฏธ์ง์ ๊ณ ์ ํ ํํ์ ์ง๊ฐ์ ํ์ค๊ฐ ์๊ฒ(photo-realistic) ๋ณด์กดํ๋ฉด์ ๊ฐ๋ฒผ์ด ์ฐ์ฐ์ผ๋ก ๋ ๋๋งํ๋ค.
์ํธ์์ฉ ๊ฐ๋ฅํ ํ์
๊ณต๊ฐ(sphere proxy): ๋ณํ๋ 3D ๋ฌผ์ฒด ์ฃผ์์ ๋ฐํฌ๋ช
ํ ๊ตฌ์ฒด ์ถฉ๋์ฒด(sphere proxy)๊ฐ ์์ฑ๋์ด, ์ฌ์ฉ์๋ ์ด๋ฅผ ์์ด๋ ์ปจํธ๋กค๋ฌ๋ก ์ก๊ณ , ํฌ๊ธฐ๋ฅผ ํค์ฐ๊ณ , ์๊ฒฉ ๋๋ฃ์ ํจ๊ป ์ด๋ฆฌ์ ๋ฆฌ ๋๋ ค๋ณด๋ฉฐ ๋ฌผ๋ฆฌ์ ๋ฐฐ์น๋ ๋์์ธ์ ์์ ๋กญ๊ฒ ๋
ผ์ํ ์ ์๋ค.
์ฐ๊ตฌ๊ฒฐ๊ณผ ๋ฐ ์์
์ปค๋ฎค๋์ผ์ด์ ํจ์จ ๊ทน๋ํ(N=12, N=18): ์ฌ์ฉ์๊ฐ ์ง์ ์์ผ๋ก 3D ๋ฌผ์ฒด๋ฅผ ์กฐ์ํ๋ฉฐ ์ค๋ช ํ ์ ์๊ฒ ๋๋ฉด์, ๋๋ฃ์๊ฒ ๋ณต์กํ ๊ณต๊ฐ์ ๊ฐ๋ ์ ์ดํด์ํค๋ ์ธ์ง์ ๋ ธ๋ ฅ์ด ํฌ๊ฒ ์ค์ด๋ค๊ณ ๋ํ ์ผํ ์๊ฐํ ํผ๋๋ฐฑ ์ฑ๋ฅ์ด ๋ํญ ํฅ์๋์๋ค.
2D์ 3D์ ๋งฅ๋ฝ๋ณ ์ํธ ๋ณด์์ฑ ๋ฐ๊ฒฌ: ์์ด๋์ด๋ฅผ ์์ ๋กญ๊ฒ ๋ฐ์ฐํ๊ณ ์กฐ์จํ๋ ์ด๊ธฐ ๋ธ๋ ์ธ์คํ ๋ฐ ๋จ๊ณ์์๋ 3D ๋ฌผ์ฒด ์กฐ์์ด ์๋์ ์ผ๋ก ์ ๋ฆฌํ์ผ๋, ์ต์ข ๊ฒฐ๊ณผ๋ฌผ์ ์์ฝ/๋ฐํํ๋ ์ ๋ฆฌ ๋จ๊ณ์์๋ ์คํ๋ ค ๊น๋ํ๊ฒ ์ ๋๋ 2D ์ค๋ ์ท ๋ทฐํฌํธ๋ฅผ ํ์ดํธ๋ณด๋์ ์ ๋ ฌํ๋ ํํ๊ฐ ์ ํธ๋๋ ๋ฑ ๋ช ํํ ์ญํ ๋ถ๋ด์ ํ์ธํ๋ค.
XR ์๊ฒฉ ์ค์ฌ๊ฐ(Copresence)์ ํ์ฅ: ๊ธฐ์กด XR ์ฐ๊ตฌ๊ฐ ์ธ๊ฐ ์๋ฐํ๋ฅผ ์ผ๋ง๋ ์ค๊ฐ ๋๊ฒ ๋ง๋๋๋์ ์ฃผ๋ก ์น์คํ๋ค๋ฉด, ๋ณธ ์ฐ๊ตฌ๋ ํ์์ค ์์ artifacts ๋ํ ์ธ๊ฐ๋งํผ ์ฆ๊ฐ์ ์ด๊ณ ์ฌ์ค์ ์ผ๋ก ๊ณต์ ๋์ด์ผ ๋น๋ก์ ์์ฑ๋ ๋์ ์๊ฒฉ ํ์
์ด ๊ฐ๋ฅํ๋ค๋ ์๋ก์ด ๋ฐฉํฅ์ฑ์ ์ ์ํ๋ค.
์ฅ์
๊ทน๋์ ์ ๋ ฅ ํธ์์ฑ: ์ฌ๋ฌ ๋ฒ ํด๋ฆญํ ํ์ ์์ด ์ปจํธ๋กค๋ฌ ํธ๋ฆฌ๊ฑฐ๋ฅผ ์ฅ๊ณ ๊ธ๋ continuous marking ์ก์ ๋ง์ผ๋ก ๋ณต์กํ ๋ ์ด์์์์ ์ํ๋ ๊ฐ์ฒด๋ฅผ ์ ํํ๊ฒ ์ถ์ถํ๋ค.
2D Pie Menu: 3D ์์ฑ ์ , ์ปจํธ๋กค๋ฌ์ ๋ถ์ฐฉ๋ 2D ํ์ด ๋ฉ๋ด๋ฅผ ํตํด 4๊ฐ ์ง๊ต ๋ทฐ(Front, Side, Back)์ 360๋ ๋น๋์ค ๋ฏธ๋ฆฌ๋ณด๊ธฐ๋ฅผ ํ๋ผ์ด๋นํ๊ฒ ๊ฒํ ํ ์ ์์ด ๊ณต์ ๊ณต๊ฐ์ ์ด์ง๋ฝํ์ง ์๋๋ค.
์๋ฐฉํฅ ๋ฏธ๋์ด ๋ณํ(3D-to-2D): ์์ฑ๋ 3D ๊ฐ์ฐ์์ ๊ฐ์ฒด๋ฅผ ์ํ๋ ๊ฐ๋์์ ์ค๋
์ท์ผ๋ก ์บก์ฒํด ๊ฐ์ ํ์ดํธ๋ณด๋์ 2D ์คํฐ์ปค์ฒ๋ผ ๋ถ์ผ ์ ์์ด ๊ฐํํ์ ํ์๊ณผ ๋ฌธ์ํ๊ฐ ๋์์ ๊ฐ๋ฅํ๋ค.
ํ๊ณ์
์ ๋ฐ ์๊ฐํ์ ํ๊ณ(blurriness): ์์ฑ ์๋์ conversational flow๋ฅผ ํ๋ณดํ๊ธฐ ์ํด ๊ฒฝ๋ ํ์ดํ๋ผ์ธ์ ์ฑํํ์ฌ, ์ฌ์ธํ ํ ์ค์ฒ ํจํด(์, ์ท์ ์คํธ๋ผ์ดํ ๋ฌด๋ฌ)์ด๋ ์๋ฃ ์ง๋จ์ฒ๋ผ ๊ณ ์ ๋ฐ๋๊ฐ ์๊ตฌ๋๋ ์ต์ข ์์ฌ๊ฒฐ์ ๋จ๊ณ์๋ 3D ๋ชจ๋ธ์ด ๋ค์ ํ๋ฆฟํ๊ฒ ๋ณด์ผ ์ ์๋ค.
์ด๊ธฐ ํฌ๊ธฐ(scale) ์์ธก์ ๋ถํ์ค์ฑ: ๋จ ํ ์ฅ์ 2D ์ด๋ฏธ์ง ์์ค์ ์์กดํ์ฌ 3D ๊ฐ์ฒด๋ฅผ ์์ฑํ๊ธฐ ๋๋ฌธ์, ์ฐธ์กฐํ ๋ฌผ๋ฆฌ์ ์นดํผ๊ฐ ์๋ ๋์งํธ ์ด๋ฏธ์ง์ ๊ฒฝ์ฐ ์ด๊ธฐ ์์ฑ ์ ์ค์ ์ฌ๋ฌผ ํฌ๊ธฐ์ ๋ค๋ฅด๊ฒ ๋ ๋๋ง๋๋ ๋ถํผ ๋ถ์ผ์น ๋ฌธ์ ๊ฐ ๋ฐ์ํ๋ค.
ํฅํ๊ณผ์
ํ ์ค์ฒ ๊ณ ๋ํ ๋ฐ PBR ์ตํฉ: ๋ ๋๋ง ์ฑ๋ฅ์ ์ ํ์ํค์ง ์๋ ์ ์์ ์ ๋ ฅ ๋ฐ์ดํฐ ๋ฐ๋๋ฅผ ๋์ด๊ณ , ๋ฌผ๋ฆฌ ๊ธฐ๋ฐ ๋ ๋๋ง(PBR) ๊ธฐ์ ์ ํตํฉํ์ฌ ๊ฐ์ ์ฌ๋ฌผ์ ํ๋ฉด ์ฌ์ง๊ฐ๊ณผ ์กฐ๋ช ๋ฐ์ฌ ํจ๊ณผ๋ฅผ ํ ์ฐจ์ ๋ ๋์ด์ฌ๋ฆด ๊ณํ์ด๋ค.
์ถ์์ ๊ฐ๋ ์ ์๊ฐํ ๋ฐ suggestion ์์คํ : ๋ฌผ๋ฆฌ์ ์ค์ฒด๊ฐ ์๋ ์ถ์์ ์ธ ์์ด๋์ด๋ ๊ฐ๋ ๋ ๋ฉํํฌ์ , ์์ง์ ์๊ฐ ๊ธฐ๋ฒ์ ์ ์ฉํด 3D assets์ผ๋ก ๊ตฌ์ฒดํํ ์ ์๋๋ก ์ปดํจํฐ ๋น์ ๊ธฐ๋ฐ์ ๋ฌธ๋งฅ ๋ง์ถคํ ์ ์ ์๊ณ ๋ฆฌ์ฆ์ ํ์ฅํ ์์ ์ด๋ค.
Adil Rahman, Rifat Rahman Khan, Jonggi Hong, Stephanie Valencia, Seongkook Heo
CustomSight: Enhancing LLM-Powered Visual Assistance for Blind Individuals using Goal-Directed Dynamic Filters
(Abstract) LLM-powered assistive technologies (ATs) have enabled blind and visually impaired (BVI) users to query personalized, goal-oriented information about their visual environment. However, the accuracy of system responses depends heavily on well-framed, queryrelevant images, which can be difficult for BVI users to capture. We present CustomSight, an LLM-powered AT that helps BVI users effectively query visual information by providing task-aware, real-time guidance to frame the camera and automatically capture images when relevant content is in view. When a user issues a query, CustomSight generates a Dynamic Filterโa custom pipeline that encodes logic tied to the userโs intent, monitors the live feed, and triggers context-aware feedback and image capture. The captured image is sent to the LLM to fetch accurate visual information.
(Introduction) Blind and visually impaired (BVI) individuals frequently rely on smartphone-based ATs to independently access visual information in daily life. Several computer vision (CV) modelsโsuch as text recognition, object detection, and scene captioningโhave been integrated into popular ATs (e.g., Seeing AI [9], Tap Tap See [5]) to help users explore their surroundings and identify objects [6, 7]. However, despite widespread adoption, these solutions typically provide context-agnostic responses, forcing BVI users to creatively โhackโ around limitations or cope with significant cognitive load to find relevant information [7].
Recently, LLM-powered ATs (e.g., Be My AI [2]) have introduced natural language querying, allowing BVI users to ask targeted visual questions conversationally for diverse use cases [1, 8, 10]. However, these systems depend critically on capturing relevant, well-framed images to produce useful responsesโan inherently difficult task without sight [3]. For instance, if a user wants to know a food itemโs expiry date, the image must clearly show it. Given the asynchronous nature of LLM responses [4], capturing a usable image may take multiple attempts, reducing efficiency and limiting practicality.
To address this, we designed CustomSight, an assistive system that creates custom pipelines to automatically capture images relevant to a userโs goal and query the LLM, helping BVI individuals access goal-oriented visual information effectively. CustomSight retains a chat-based interface similar to existing LLM-powered apps like Be My AI. However, instead of simply facilitating question answering, CustomSight first checks whether the submitted image contains enough information to meaningfully answer the query. If not, it assesses whether available on-device sensing and feedback modules can help the user better aim the camera and trigger capture when relevant content is likely in frame. It then generates a Dynamic Filterโa tailored pipeline that processes the live feed using goal-specific logic, provides context-aware audio cues to assist framing, and automatically captures an image once the userโs goal criteria are met, passing it to the LLM for a task-specific response (Fig 2a). For example, if a user looks for expiry date, CustomSight may generate a Dynamic Filter that monitors text recognition output for relevant keywords like โexpiry dateโ, โuse byโ, and โbest beforeโ, provides continuous sonification and speech feedback, and captures when relevant content is detected and steady (Fig 2b).
(Conclusion) We presented the design of CustomSight, an LLM-powered visual assistive tool that helps blind users efficiently capture task-relevant images through real-time, goal-directed guidance. Our preliminary study highlighted its potential to improve the efficiency of visual information access for BVI users. Future work includes co-design workshops to explore new sensing and feedback modules (e.g., depth, haptics), supporting more expressive framing guidance, and real-world deployments to assess everyday utility and impact.
๊ธฐ์กด ์๊ฐ ๋ณด์กฐ ๊ธฐ์ (AT)์ ๋ณํ์ ํ๊ณ
์ปดํจํฐ ๋น์ ๊ธฐ๋ฐ ์ฑ(์, Seeing AI ๋ฑ): ํ ์คํธ ์ธ์(OCR)์ด๋ ๋จ์ ๋ฌผ์ฒด ๊ฐ์ง ๊ธฐ๋ฅ์ ์ ๊ณตํ์ง๋ง, ์ฌ์ฉ์์ ๊ตฌ์ฒด์ ์ธ ๋ชฉ์ ์ด๋ ๋งฅ๋ฝ์ ๊ณ ๋ คํ์ง ์์ ์ผ๋ฐฉํฅ์ ์ด๊ณ ๊ตฌ์กฐํ๋์ง ์์ ์ ๋ณด(context-agnostic)๋ง ๋์ดํ๋ค. ์ด๋ก ์ธํด ์๊ฐ์ฅ์ ์ธ ์ฌ์ฉ์๊ฐ ์ ์ ์ํ๋ ์ ๋ณด๋ฅผ ์ฐพ์ผ๋ ค๋ฉด ์๋นํ ์ธ์ง์ ๋ถ๋ด์ ์ง๊ฑฐ๋ ์์คํ ์ ์ฐํํ๋ ํธ๋ฒ์ ๋์ํด์ผ ํ๋ค.
์ต๊ทผ LLM ๊ธฐ๋ฐ ๋ํํ ์ฑ(์, Be My AI ๋ฑ): ์์ฐ์ด ์ง๋ฌธ(VQA, Visual Question Answering)์ ํตํด ์ฌ์ฉ์ ๋ง์ถคํ ๋ต๋ณ์ ์์ธํ๊ฒ ์ ๊ณตํ ์ ์๊ฒ ๋์์ผ๋, ์ด ์์คํ ๋ค์ ์ง๋ฌธ ๋ชฉ์ ์ ๋ง๊ฒ ์ ์ดฌ์๋ ์ด๋ฏธ์ง๊ฐ ์ ์ ๋์ด์ผ๋ง ์ ํํ ๋ต๋ณ์ ์ค ์ ์๋ค๋ ๋จ์ ์ด ์๋ค.
ํ์ค์ ์ธ ์ดฌ์์ ์ด๋ ค์: ์๋ ฅ์ด ์๋ ์ํ์์ ํน์ ์ ๋ณด(์, ์ํ์ ์ ํต๊ธฐํ ์์น)๊ฐ ์ ํํ ์ฐํ๋๋ก ์นด๋ฉ๋ผ ์ต๊ธ์ ์ก๋ ๊ฒ์ ๋งค์ฐ ์ด๋ ต๋ค. ๊ฒ๋ค๊ฐ LLM์ ๋ต๋ณ์ ๋น๋๊ธฐ์์ผ๋ก ๋๋ฆฌ๊ฒ ์ค๊ธฐ ๋๋ฌธ์, ์ ๋๋ก ๋ ์ฌ์ง์ด ์ฐํ ๋๊น์ง ์์ฐจ๋ก ์ดฌ์๊ณผ ๋๊ธฐ๋ฅผ ๋ฐ๋ณตํด์ผ ํ๋ฏ๋ก ํจ์จ์ฑ์ด ํฌ๊ฒ ๋จ์ด์ง๋ค.
CustomSight
์ฌ์ฉ์๊ฐ ์์ฐ์ด๋ก ์ง๋ฌธ์ ๋์ง๋ฉด AI๊ฐ ๊ทธ ์๋๋ฅผ ํ์ ํด ๋ง์ถคํ ์นด๋ฉ๋ผ ํํฐ(dynamic filter)๋ฅผ ์ฆ์์์ ์์ฑํ๊ณ , ์ค์๊ฐ ์จ๋๋ฐ์ด์ค ์ผ์ฑ ๋ฐ ์ค๋์ค ํผ๋๋ฐฑ์ ํตํด ์ฌ๋ฐ๋ฅธ ์ดฌ์์ ์ ๋ํ ๋ค ๋ชฉ์ ์ ๋ง๋ ํ๋ฉด์ด ํฌ์ฐฉ๋๋ฉด ์๋์ผ๋ก ์ฌ์ง์ ์บก์ฒํ๋ ํ์ด๋ธ๋ฆฌ๋ ์๊ฐ ๋ณด์กฐ ์์คํ ์ด๋ค.
๋ต๋ณ ๊ฐ๋ฅ์ฑ ํ๋ณ(determine answerability): ์ฌ์ฉ์๊ฐ ์ฒ์ ์ฌ์ง๊ณผ ์ง๋ฌธ์ ๋ณด๋์ ๋, ํ์ฌ ์ฌ์ง๋ง์ผ๋ก ๋ต๋ณ์ด ๊ฐ๋ฅํ์ง LLM์ด rubric ๊ธฐ๋ฐ ํ๋กฌํํธ๋ก ์ฌ์ฌํ๋ค. ์ ๋ณด๊ฐ ๋ถ์กฑํ๋ค๊ณ ํ๋จ๋๋ฉด ์ฌ์ดฌ์(retake) ๋จ๊ณ๋ก ์ง์ ํ๋ค.
ํํฐ ์์ฑ ๊ฐ๋ฅ์ฑ ํ๊ฐ(determine filter programmability): ์ค๋งํธํฐ์ ์จ๋๋ฐ์ด์ค ์ผ์ ๋ชจ๋(Apple VisionKit ํ ์คํธ ์ธ์, YOLOv8 ๋ฌผ์ฒด ๊ฐ์ง ๋ฑ)๊ณผ ์ค๋์ค ํผ๋๋ฐฑ ๋ชจ๋์ ์กฐํฉํด ์ดฌ์์ ๋์ธ ์ ์๋ ๊ฐ์ด๋ ๋ก์ง์ ์งค ์ ์๋์ง LLM์ด ํ๋จํ๋ค.
๋์ ํํฐ ์์ฑ(generate dynamic filter): ๊ฒ์ฆ์ ํต๊ณผํ๋ฉด LLM์ด ์ ์ฉ ๋๋ฉ์ธ ํนํ ์ธ์ด(DSL, Domain-Specific Language)๋ฅผ ์ฌ์ฉํ์ฌ ์ฌ์ฉ์ ๋ชฉ์ ์ ๋ง์ถ ๊ฐ์ง ์๊ณ ๋ฆฌ์ฆ์ ์ค์๊ฐ์ผ๋ก ์ค๊ณํ๋ค, e.g., "์ ํต๊ธฐํ ์ธ์ ์ผ?"๋ผ๊ณ ๋ฌผ์ผ๋ฉด, OCR ๋ชจ๋์ด 'best before', 'use by', 'exp' ๊ฐ์ ๋จ์ด๋ ๋ ์ง ์ ๊ท์ ํจํด์ ๊ฐ์งํ๋๋ก ํํฐ๋ฅผ ๊ตฌ์ฑํ๋ค.
ํํฐ ์คํ ๋ฐ ์๋ ์บก์ฒ(execute filter): ์ฌ์ฉ์๊ฐ ์นด๋ฉ๋ผ๋ฅผ ์์ง์ด๋ฉด ์์คํ
์ด ์ค์๊ฐ ๋น๋์ค ํ๋ ์์ ๋ชจ๋ํฐ๋งํ๋ค. ์ํ๋ ์ ๋ณด๊ฐ ํ๋ฉด์ ๋ค์ด์ค๊ณ ์๋จ๋ฆผ ์์ด ์์ ๋๋ ์๊ฐ(stable), ์ค๋์ค ์ ํธ(tone ๋ชจ๋ ๋ฐ TTS)๋ก ๊ฐ์ด๋๋ฅผ ์ฃผ๋ฉฐ ์ต์ ์ ํ ์ปท์ ์์์ ์๋ ์บก์ฒ(auto-capture)ํ์ฌ LLM์ ์ ์กํ๋ค.
์ฐ๊ตฌ๊ฒฐ๊ณผ ๋ฐ ์์
์ํธ์์ฉ์ฑ ๋ฐ ํจ์จ์ฑ ํฅ์: ๋ ๋ช ์ ์๊ฐ์ฅ์ ์ธ ์ฐธ๊ฐ์(์ ๋ฌธ AT ์คํ์ ๋ฆฌ์คํธ ๋ฐ ์ค๋ ์ค๋ช ์)๋ฅผ ๋์์ผ๋ก 4๊ฐ์ง ์ผ์ ์์ (์๋ฆฌ๋ฒ ์ฐพ๊ธฐ, ํต์กฐ๋ฆผ ์ ํต๊ธฐํ ํ์ธ, ์ค๋ ์ง ์ฐพ๊ธฐ, ๋ฉ๋ดํ ๊ฐ๊ฒฉ ํ์ธ)์ ์คํํ ๊ฒฐ๊ณผ, ๊ธฐ์กด ๋ํํ ์ฑ๋ณด๋ค ํจ์ฌ ๋ ์ ์ ๋จ๊ณ๋ง์ผ๋ก ์ํ๋ ์ ๋ณด๋ฅผ ์ ํํ๊ณ ํจ์จ์ ์ผ๋ก ํ๋ํ๋ค.
์์คํ ์ ๋ํ ์ ๋ขฐ๋(trust) ์์น: ์ฌ์ฉ์๊ฐ ์ธ์ ์ ํฐ๋ฅผ ๋๋ฅผ์ง ๋ถ์ํดํ ํ์ ์์ด, ์์คํ ์ด ์ ์ ํ ํ๋ ์์ ์ค์ค๋ก ํ๋จํ๊ณ ํ์ ์ ๊ฐ์ง๋ฉฐ ์๋์ผ๋ก ์ฐ์ด์ค๋ค๋ ์ ์ด ์ฌ์ฉ์์๊ฒ ์ฌ๋ฆฌ์ ์์ ๊ฐ๊ณผ ๋์ ์ ๋ขฐ๋ฅผ ์ ๊ณตํ์ต๋๋ค.
ํ์ด๋ธ๋ฆฌ๋ ํจ๋ฌ๋ค์ ์ ์: ๋ณธ ์ฐ๊ตฌ๋ ์ค์๊ฐ ์จ๋๋ฐ์ด์ค ์ปดํจํฐ ๋น์ (CV/ML)์ ์๋๊ฐ๊ณผ ๊ฑฐ๋ ์ธ์ด ๋ชจ๋ธ(LLM)์ ๊ณ ์ฐจ์์ ๋ฌธ๋งฅ ์ดํด๋ ฅ์ ๊ฒฐํฉํ์ฌ, ๋ ๊ธฐ์ ์ ํ๊ณ๋ฅผ ๋์์ ๊ทน๋ณตํ๋ ์ฐจ์ธ๋ ์ ๊ทผ ๋ฐฉ์์ ์ ์ํ๋ค๋ ์ ์์ ํ์ ์ ์์๊ฐ ํฝ๋๋ค.
์ฅ์
๋ชฉ์ ์งํฅํ ๊ฐ์ด๋ฉ: ์ฌ์ฉ์ ์ง๋ฌธ์ ๋ฐ๋ผ ๊ทธ๋๊ทธ๋ DSL ์ฝ๋๊ฐ ์์ฑ๋๋ฏ๋ก, ๊ณ ์ ๋ ๊ธฐ๋ฅ๋ง ์ ๊ณตํ๋ ์ผ๋ฐ ์๊ฐ ์ฑ๊ณผ ๋ฌ๋ฆฌ ๋ฌดํํ ์ผ์ ์๋๋ฆฌ์ค์ ์ ์ฐํ๊ฒ ๋์ํฉ๋๋ค.
์ค์๊ฐ on-device ์ฒ๋ฆฌ: React Native ๊ธฐ๋ฐ ์ฑ๊ณผ Flask/Socket.IO ๋ฐฑ์๋๋ฅผ ์ฐ๋ํ์ฌ ๋น๋์ค ํผ๋๋ฅผ ๋๊น ์์ด ์ฒ๋ฆฌํ๊ณ ์ค์๊ฐ ์ค๋์ค ํผ๋๋ฐฑ์ ์ ๋ฌํฉ๋๋ค.
ํ๊ณ์
๊ฐ์ฅ์๋ฆฌ ์กฐ๊ธฐ ์บก์ฒ ํ์(early trigger): ํ์ฌ ์์คํ
์ ํค์๋ ํจํด์ ์ ๋ฌด(binary inclusion) ์์ฃผ๋ก ํ๋จํ๊ธฐ ๋๋ฌธ์, 'cooking instructions'๋ผ๋ ๊ธ์๊ฐ ํ๋ฉด ๊ตฌ์ ๊ฐ์ฅ์๋ฆฌ์ ์ด์ง๋ง ๊ฑธ์ณ๋ ๋๋ฌด ๋นจ๋ฆฌ ์บก์ฒ๋ฅผ ํด๋ฒ๋ ค ์ ์ ์๋งน์ด ๋ด์ฉ์ด ์ฌ์ง์์ ์๋ฆฌ๋ ํ๊ณ๊ฐ ์์ต๋๋ค.
ํฅํ๊ณผ์
๋ฏธ์ธ ์ ์ด ํผ๋๋ฐฑ ๊ณ ๋ํ: ํ๊น ์ ๋ณด๊ฐ ํ๋ฉด ์ ์ค์(center)์ ์์ ์ ์ผ๋ก ์์นํ ์ ์๋๋ก ์ ๋ํ๋ ์ ๋ฐํ ์ค๋์ค ๋ฐฉํฅ ๊ฐ์ด๋(sonification) ๊ธฐ์ ์ด๋ ์ค๋งํธํฐ ์ง๋(haptics) ํผ๋๋ฐฑ ์๊ณ ๋ฆฌ์ฆ์ ์ถ๊ฐํ ์์ ์ด๋ค.
๋ฉํฐ๋ชจ๋ฌ ์ผ์ ๋ฐ ํ IoT ํ์ฅ: depth ์นด๋ฉ๋ผ ์ผ์๋ฅผ ์ตํฉํด ๋ณดํ ์ค ์ถฉ๋ ๋ฐฉ์ง ๊ธฐ๋ฅ์ ๊ตฌํํ๊ฑฐ๋, ์ค๋งํธ ํ ์นด๋ฉ๋ผ์ ์ฐ๋ํ์ฌ ํ๊ด ์ ํ๋ฐฐ ๋ฌผํ์ ๊ฐ์งํ๋ ๋ฑ ์ผ์ ๋ฐ์ฐฉํ ๋ชจ๋์ ํ์ฅํ๊ณ ์ค์ ํ๊ฒฝ์์ ์ฅ๊ธฐ ๋ฐฐํฌ ์ฐ๊ตฌ๋ฅผ ์งํํ ๊ณํ์ด๋ค.ย
Hyeongjin Kim, Erzhen Hu, Seongkook Heo
SpaceShare: Leveraging Multimodal Context for Fluid Sharing of Spaces in Video Meetings
(Abstract) In video meetings, people not only talk but also share objects and their environments. However, conventional video calls offer limited support for spatial sharing, typically relying on a live video feed. We present SpaceShare, a system that integrates real-time 3D space reconstruction into video meetings and leverages multimodal context to fluidly support shared spatial understanding. The reconstructed 3D environment enables users to independently explore the space, while avatar visualization and shared views convey each participantโs location and viewpoints. SpaceShare also stores conversation content and contextual metadata, such as where the conversation took place. This allows users to retrieve spatial features using both physical attributes and conversational context.
(Introduction) Video meetings are now widely used across many areas. Beyond simple conversation, people increasingly share physical environments through video meetings. For example, they take virtual guided tours [6], view real estate remotely [4], and collaborate on spatial projects [1]. While modern mobile devices can capture scenes in high fidelity, effectively communicating spatial information remains challenging. The video feed is momentary and shows only what is currently in view. Referencing parts of the space is difficult and inefficient, especially when the location is out of view or far from the users. As the shared space becomes larger, problems become more pronounced.
To effectively support space sharing in video meetings, prior work has identified several key aspects: enabling view independence, fostering spatial awareness, and supporting referencing of past scenes. View independence allows users to explore remote environments freely, improving task efficiency and confidence [15], though solutions like 360ยฐ video remain limited by fixed viewpoints [1, 16, 17]. Real-time 3D reconstruction using Simultaneous localization and mapping (SLAM) methods [14] or neural techniques like NeRF and Gaussian Splatting [7, 12] may address this by allowing interactive, live exploration. For spatial awareness, systems have used avatars, gaze, and gestures to convey partner context [2, 10], but these become less effective in large spaces. Finally, referencing previously visited scenes remains challenging. Landmarks and minimaps help [18], but are often insufficient without integration of conversational context [9].
To address these challenges, we developed SpaceShare, a system that integrates real-time 3D space reconstruction with multimodal contextual features to enhance shared spatial understanding in video meetings. Using SLAM, SpaceShare reconstructs the environment on the fly, enabling both local and remote users to independently explore the shared space. The system visualizes each userโs position and viewpoint through virtual avatars and ray-based pointers to facilitate view communication. As users converse, SpaceShare captures not only the 3D environment but also multimodal contextual data, including user locations, and conversation content. It generates panoramic images by stitching keyframes, labels objects in the scene, and stores these representations alongside the conversation history. This enables rich, context-based retrieval of prior scenes. For instance, instead of navigating manually or issuing a generic search like โchair,โ a user could search for โthe chair that reminded us of our childhood memories,โ leveraging the conversational context to find a specific moment or location.
Our key contribution is the introduction of a novel approach that incorporates multimodal contextual summaries to support spatial sharing and retrieval in remote meetings.
(Conclusion) While our preliminary study demonstrated the potential of SpaceShare for enhancing spatial recall and collaboration in remote meetings, the SUS scores indicate that the overall interface and interaction design still require improvement. Future work will include iterative design refinements to improve usability and a more comprehensive user study to gain deeper insights into the effects of the space summarization and retrieval features.
๊ธฐ์กด ์์ ๊ณต์ ๋ฐฉ์์ ํ๊ณ
์๊ฐ์ ๋จ์ (momentary video feed): ์ผ๋ฐ ํ์ ํ์๋ ์นด๋ฉ๋ผ๊ฐ ๋น์ถ๋ ์๊ฐ์ 2D ํ๋ฉด๋ง ์ ํ์ ์ผ๋ก ๋ณด์ฌ์ฃผ๋ฏ๋ก, ํ๋ฉด ๋ฐ์ ์ฃผ๋ณ ๊ณต๊ฐ์ ์ค๋ช ํ๊ฑฐ๋ ๋ฉ๋ฆฌ ๋จ์ด์ง ๊ณต๊ฐ์ ํน์ ์์ญ์ referencingํ๊ธฐ๊ฐ ๊ทน๋๋ก ๋ถํธํ๊ณ ๋นํจ์จ์ ์ด๋ค.
360ยฐ ๋น๋์ค์ ํ๊ณ(fixed viewpoints): ์๊ฒฉ ์ฌ์ฉ์๊ฐ ํ๋ฉด์ ๋๋ ค๋ณผ ์๋ ์์ง๋ง, ์นด๋ฉ๋ผ๊ฐ ๋์ธ ๊ณ ์ ๋ ๋ฌผ๋ฆฌ์ ์์น(fixed perspective)์์๋ง ๋ฐ๋ผ๋ด์ผ ํ๋ฏ๋ก ์์ ๋ก์ด ๊ณต๊ฐ ํ์(view independence)์ด ๋ถ๊ฐ๋ฅํ๋ค.
๊ณผ๊ฑฐ ๋งฅ๋ฝ ์์ค(loss of context): ์ด์ ์ ๋ฐฉ๋ฌธํ๊ฑฐ๋ ๋๋ด๋ ๋ํ ๋ด์ฉ์ด ๊ณต๊ฐ ์ ๋ณด์ ์ ๊ธฐ์ ์ผ๋ก ๊ฒฐํฉ๋์ง ์๋๋ค, e.g., ํ์์ค์ "์๊น ๋ํํ๋ ๊ฑฐ๊ธฐ ์ด๋์์ง?"๋ผ๋ฉด, ์ด์ ์ฅ๋ฉด์ ๋ฌผ๋ฆฌ์ ์์น๋ฅผ ๋ค์ ์ถ์ ํ๊ธฐ๊ฐ ๋งค์ฐ ๋ํดํ๋ค.
SpaceShare
ํ์ฅ(local) ์ฌ์ฉ์๊ฐ ํ๋ธ๋ฆฟ์ด๋ ์นด๋ฉ๋ผ๋ก ์ฃผ๋ณ์ ๋น์ถ๋ฉด ์ค์๊ฐ์ผ๋ก 3D ๊ณต๊ฐ์ ๋ณต์ํ๊ณ , ์ฐธ๊ฐ์๋ค์ ๋ํ ๋ด์ฉ(transcript)๊ณผ 3D ์์น ์ขํ๋ฅผ ๊ฒฐํฉํ์ฌ "๊ทธ๋ ๊ฑฐ๊ธฐ์ ๋๋ ๋ํ"์ ๋งฅ๋ฝ์ผ๋ก ํน์ ์ฅ์๋ฅผ ์ญ์ถ์ ํด๋ด๋ ํ์ ์ ์ธ ํ์ ํ์ ํ๋ซํผ์ด๋ค.
์ค์๊ฐ 3D ๊ณต๊ฐ ๋ณต์(SLAM ๊ธฐ๋ฐ): Spectacular AI Visual-Inertial SLAM SDK๋ฅผ ํ์ฉํด depth ์นด๋ฉ๋ผ ์ผ์ ๋ฐ์ดํฐ๋ก๋ถํฐ ์ค์๊ฐ 3D ๋ฉ์ฌ์ ํฌ์ธํธ ํด๋ผ์ฐ๋๋ฅผ ์ถ์ถํ๋ค. ์๊ฒฉ ์ฌ์ฉ์๋ ๊ฐ์ ์๋ฐํ์ ray-based pointers๋ฅผ ํตํด ๋ฐฉ์ก ํ๋ฉด์ ๊ฐํ์ง ์๊ณ ๋ ๋ฆฝ์ ์ผ๋ก 3D ๊ณต๊ฐ์ ๋์๋ค๋ ์ ์๋ค.
๋ฉํฐ๋ชจ๋ฌ ๊ณต๊ฐ ์์ฝ(multimodal space summarization): ์นด๋ฉ๋ผ ํ๋ ์์ ์๊ฐ์ ์ฐจ์ด๋ฅผ ๊ฐ์งํด ํต์ฌ ์ฅ๋ฉด์ ์ถ์ถํ ๋ค, OpenCV๋ก ํ๋ ธ๋ผ๋ง ์ด๋ฏธ์ง๋ฅผ ๋ณํฉํ๋ค. ์ฌ๊ธฐ์ YOLO-World๋ก ์ฌ๋ฌผ์ ์๋ ๋ผ๋ฒจ๋งํ๊ณ OpenAI GPT-4V๋ก ํ ์คํธ ์ค๋ช ์ ์์ฑํ๋ ๋์์, OpenAI Speech-to-Text API๋ก ๋ํ ๋ด์ฉ์ ํ์์คํฌํ์ ๊ฒฐํฉํด ํ๋์ ๋ฉ๋ชจ๋ฆฌ ๋ฐ์ดํฐ๋ฒ ์ด์ค๋ก ๊ตฌ์ถํ๋ค.
๋งฅ๋ฝ ๊ธฐ๋ฐ ๊ณต๊ฐ ๊ฒ์(context-based space retrieval): ๋จ์ํ "์์" ๊ฐ์ ๋ฌผ๋ฆฌ์ ๋จ์ด๋ฅผ ๊ฒ์ํ๋ ๋์ , OpenAI GPT-4o๋ฅผ ํตํด "์ฐ๋ฆฌ ์๊น ์ด๋ฆด ์ ๊ธฐ์ต ์๊ธฐํ๋ฉด์ ์ฅ๋์ณค๋ ๊ทธ ์์"์ฒ๋ผ ๋ํ ์ ์ํผ์๋ ๋งฅ๋ฝ์ ์์ฐ์ด๋ก ์
๋ ฅํ๋ฉด ์์คํ
์ด ์ ํํ 3D ์์น ์ขํ์ ๋น์์ ํ๋
ธ๋ผ๋ง ์ฌ์ง์ ์ฆ๊ฐ ์ฐพ์๋ธ๋ค.
์ฐ๊ตฌ ๊ฒฐ๊ณผ ๋ฐ ์์
๊ณต๊ฐ ํ์ ๋ฅ๋ ฅ ํฅ์(exploratory study, N=16): 16๋ช ์ ์ฐธ๊ฐ์๋ฅผ ๋์์ผ๋ก ๊ณผ๊ฑฐ์ ๋ํ ๋๋ด๋ ํน์ ์์ ํ 5๊ฐ๋ฅผ ๋ค์ ์ฐพ์๋ด๋ ๋ฏธ์ ์ ์ํํ ๊ฒฐ๊ณผ, ๋งฅ๋ฝ ๊ธฐ๋ฐ ๊ณต๊ฐ ๊ฒ์ ๊ธฐ๋ฅ์ ์ฌ์ฉํ์ ๋(ํ๊ท 4.44๊ฐ ์ฑ๊ณต)๊ฐ ํด๋น ๊ธฐ๋ฅ ์์ด ์๋์ผ๋ก ์ฐพ์์ ๋(ํ๊ท 3.62๊ฐ ์ฑ๊ณต)๋ณด๋ค ํจ์ฌ ๋ ๋ง๊ณ ์ ํํ๊ฒ ์์ ํ์ ์์น๋ฅผ ๋์ถํด ๋๋ค.
UI/UX ๊ฐ์ ์ ํ์์ฑ ๋ฐ๊ฒฌ: ๊ธฐ๋ฅ์ ์ ์ฉ์ฑ๊ณผ ํจ์จ์ฑ์ ์คํ์ ํตํด ๋ช ํํ ํ์ธ๋์์ผ๋, ์ฌ์ฉ์ ํธ์์ฑ ์งํ์ธ ์์คํ ์ฌ์ฉ์ฑ ํ๊ฐ(SUS ์ ์, System Usability Scale)๋ ๊ฒ์ ๊ธฐ๋ฅ์ ์ผ์ ๋ 55.47์ , ์ฐ์ง ์์์ ๋ 53.91์ ์ผ๋ก ๋ชจ๋ 50์ ๋(marginal level)๋ฅผ ๊ธฐ๋กํ๋ค.
ํฅํ ๊ณผ์ (์ธํฐํ์ด์ค ์ต์ ํ): ๋ค๋์ ๋ฉํฐ๋ชจ๋ฌ ์ ๋ณด(3D ๊ณต๊ฐ ๋ฐ์ดํฐ, ์๋ฐํ ํฌ์ธํฐ, ๋ํ ํ ์คํธ, ํ์ ์ด๋ฏธ์ง ๋ฑ)๊ฐ ํ ๋ฒ์ ์์์ง ๋ ์ฌ์ฉ์๊ฐ ๋๋ผ๋ ์๊ฐ์ ํผ๋ก๊ฐ๊ณผ ๋ณต์ก์ฑ์ ์ค์ด๊ธฐ ์ํด, ํฅํ ์ธํฐํ์ด์ค ๋์์ธ ๋ฐ ์ํธ์์ฉ ๋ ์ด์์์ ์ ๋ฉด ๋ณด์ํ ๊ณํ์ด๋ค.
Zackary T. Landsman, Matthew Clark, Seongkook Heo, Afsaneh Doryab, Gregory J. Gerling
Evaluating the Physics and Repeatability of Human Brushers in Delivering Affective Touch
(Abstract) Research in affective touch often utilizes a pleasant stroking paradigm, with touch typically delivered at velocities between 0.3โ30 cm/s and forces about 0.4 N. These values are derived from perceptual and physiological responses of touch receivers rather than natural behaviors of touchers. Herein, we observe untrained touchers in delivering pleasant touch as gentle strokes to a receiverโs forearm through high-resolution force and position measurement using an instrumented brush. Twenty participants delivered about eleven strokes in each of five trials, yielding data on stroke velocity, force, duration, and length. The cohort deployed forces (0.37 N ยฑ 0.24, 2 SD) near 0.4 N, but stroking velocities (13.7 cm/s ยฑ 8.8, 2 SD) slightly higher than typically considered โoptimalโ (1โ10 cm/s). We observed no correlation between stroke force and velocity, which led us to consider these factors jointly in characterizing an individualโs brushing strategy. Moreover, while the cohort of participants exhibited a compact range of forces and velocities, individuals tended to occupy only subsets of this range with high repeatability across trials. Altogether, the findings suggest a need to further evaluate the velocity range between 10โ30 cm/s and to jointly consider force, velocity, and consistency in characterizing an individualโs brushing strategy.
(Introduction) Affective touch typically refers to slow moving, low-force mechanical stimuli, exemplified by gestures such as a gentle caress or comforting stroke that carry emotional and social meaning, extending beyond the mere detection of physical properties of stimuli [1], [2], [3]. Affective touch often, but not always, overlaps with social touch, which occurs in interpersonal contexts and is influenced by relational factors including familiarity, intimacy, and cultural norms [4], [5]. Research into touch receiversโ perceptions of affective touch has sought to determine what makes a touch โpleasantโ and has considered inherent traits between participants, including relationships [4], [6], [7], [8], the touch receiverโs sex [9], age [10], [11], culture [5], [12], skin site [13], [14], and touch delivery via person or robot [13], [15]. In an effort to understand what makes affective touch pleasant, researchers have hypothesized that alignment between the intent of touch delivery and its interpretation by a touch recipient plays a large role [2]. Thus, to understand how the intent of touch delivery is communicated, works have evaluated how cues at physical contactโsuch as force, velocity, temperature, contact area, duration, and frequency [16], [17], [18], [19]โinfluence social and emotional perception [20] and evoke neural responses [21], [22]. Notably, even slow, gentle stroking delivered by a stranger has been shown to elicit predominantly positive affect, suggesting a potential universal language of touch that operates independently of contextual or relational factors traditionally considered essential to affective touch [23].
Pleasant touch, the perception of received touch as positive affect, sits at the intersection of affective and social touch. Studies have shown that pleasant touch can enhance physiological regulation, which contributes to stress reduction, interpersonal bonding, and social wellbeing [24]. The delivery of pleasant touch has received particular attention, as it is thought to engage a specialized class of unmyelinated low-threshold mechanoreceptors in hairy skin, C-tactile (CT) afferents, that respond optimally to gentle, caress-like stroking [3], [19]. A common experimental paradigm to research factors contributing to pleasant touch involves stroking the skin with a soft brush [13], [14], [15], [19], [20], [21], [25], [26], [27], [28], [29], [30], [31], [32], [33], [34]. As a substitute to skin-to-skin stroking, the use of a brush can remove variability in skin softness, hand size, and temperature. The brush also more readily affords instrumented sensing of force, position, and velocity, which are challenging to measure in skin-to-skin touch [15], [35]. Although stroking another person with a brush may not entirely reflect the way people touch one another in natural contexts, this controlled approach has been invaluable in studying pleasant touch.
One of the most influential findings from this line of research has been the identification of an inverted U-shaped relationship between stroking velocities and pleasantness ratings, which mirrors the firing profile of CT afferents, with peak responses in the range of ~1 โ 10 cm/s and forces in the range of 0.2 โ 0.4 N [19]. The parameters of force and velocity have become a de facto standard for delivering pleasant touch, with an โoptimal rangeโ of 1โ10 cm/s at about 0.4 N.
Force and velocity have been emphasized because they are both salient and intuitive to receivers. Touch can be readily described as โfastโ versus โslow,โ or โhardโ versus โsoft.โ For the touch deliverer, they also represent straightforward adjustments, i.e., to change velocity, one moves faster or slower; to change force, one presses harder or softer. This bidirectional clarity has contributed to their prominence in both perceptual and mechanistic studies of affective touch. Accordingly, the present study also focuses on these features. Yet critically, the six original experimental velocities (0.1, 0.3, 1, 3, 10, 30 cm/s) were selected arbitrarily to reflect โslow,โ โaverage,โ and โfastโ stroking with a brush, rather than being empirically derived from naturalistic touch behavior [19]. Over time, these same six velocities have become the default in most stroking studies [18], [25], [27], [28], [31], with an emphasis on the โoptimal rangeโ of 1 โ 10 cm/s. That said, velocities outside this range have not necessarily been deemed unpleasant, just less pleasant comparatively. Relying on these six velocities may foster a somewhat incomplete picture of how pleasant touch is naturally produced and experienced. Because these values originated from early experimental design choices rather than observations of real interactions, they may not capture the full range of motions people chose to deliver as pleasant.
In addition to the possibility of not fully capturing the preferred delivery of pleasant touch, these six velocities may mischaracterize our articulation of an โoptimal rangeโ, given notable limitations recently demonstrated at the level of individual participants. In particular, a review of five stroking studies found that the characteristic U-shaped relationship between velocity and pleasantness in the โoptimal rangeโ held for only 40% of individuals, revealing the masking effect of aggregate statistics [26]. This raises questions about which velocity ranges should guide experimental paradigms, given that some studies report positive pleasantness ratings even at 50 cm/s [36]. Advancing studies of touch delivery requires complementing perceptual research with observations of how untrained individuals naturally produce stroking gestures. Convergence between the kinematics of natural stroking patterns and pleasantness in an โoptimal rangeโ would provide evidence of receiverโdeliverer attunement, a defining characteristic of affective touch.
Prior attempts to observe natural stroking behaviors have produced mixed results. For instance, some studies observing the velocity of untrained stroke delivery report velocities exceeding the CT-optimal range (e.g., 13.9 ยฑ 4.0 cm/s, [37] and 3.9 โ 27.7 cm/s [38]), whereas others more closely align [39]. Indeed, systematic observation of stroking-based delivery is necessary, not only to assess whether its dynamics resemble or diverge from skin-to-skin stroking, but also to evaluate whether the six commonly employed experimental velocities are naturally preferred by touch deliverers.
A related issue concerns the lack of standardization in brush-based stroking delivery paradigms. Existing studies differ in several facets related to brush type, whether strokes are delivered by humans [8], [15], [20], [21], [29], [32], [33], [34] or robots [13], [14], [19], [25], [27], [28], [32], how onset and offset of skin contact are specified, and how velocity is calculated. In many cases, human experimenters are asked to approximate stroking velocity with the aid of a metronome [33], [34], or a visual cue [29], and to practice force control on a scale [15]. While each method offers a degree of calibration, the precision and consistency of the resulting strokes have not been empirically validated. Without such validation, it remains unclear to what extent the delivery of experimental stroking is consistent and reproducible. Direct measurement of the physics of stroke delivery is essential to ground laboratory approaches in objective, naturalistic observation, while also offering a framework for experimental validation of stroking delivery.
The present study addresses this gap by quantifying the force and velocity of pleasant stroking as delivered by untrained individuals with a brush. Specifically, we examine these parameters at both group and individual levels, compare them against canonical ranges, and evaluate individual consistency across repeated strokes. By capturing how people naturally apply brush strokes within a controlled setup, this study aims to provide an observational foundation for naturalistically valid, yet reproducible, models of pleasant touch delivery.
(Conclusion) This study addresses a need to describe the characteristics of untrained human brushers in delivering affective touch with a soft brush. Through direct observation in the delivery of force velocity, brushstroke length, and duration, we examined characteristics of the aggregate population and individual brushers. The observed force range aligns with that used in prior stroking studies, while the velocity range is slightly higher than the โoptimal rangeโ of 1 โ 10 cm/s. This result suggests that future stroking studies may need to evaluate velocities at a higher resolution between 10 โ 30 cm/s. Within brushers, we observed a high level of consistency in their deployment of force and velocity, which suggests that individuals employ distinctive brushing strategies.
Observations of higher brushstroke velocities
The cohort of 20 participants deployed stroke velocity with a mean of 13.38 cm/s and range of 4.87 โ 22.70 cm/s, which is somewhat higher than the typically considered โoptimal rangeโ of 1 โ 10 cm/s [14], [19], [25], [27], [32]. Likewise, other recent studies focused upon skin-to-skin contact with the hand have also reported higher ranges of velocity [37], [38]. For example, one study observing gentle stroking differences based on a receiverโs relationship to the touch deliverer reported a velocity range almost identical to that identified herein, 3.9 โ 27.7 cm/s [38]. Likewise, a mean velocity of 13.9 cm/s was observed between mothers and their 1 โ 12 month old infants [37]. It is important to note that most of this prior work evaluated the velocity of stroking in skin-to-skin touch, with comparatively little observation of velocity in brush-delivered touch. Despite the differences in delivery method, the range of velocities observed resemble one another [37], [38].
That said, with the agreement in velocities exceeding the โoptimal rangeโ of 1โ10 cm/s, a finer-grained investigation of pleasantness at higher velocities (10โ30 cm/s) is warranted. Notably, in the present study, the highest mean velocity recorded by an individual brusher was 20.33 cm/s, which was still rated as pleasant. This finding aligns with the idea that the range of 1 โ 10 cm/s represents a subset, rather than an exclusive, range of pleasantness. A study utilizing brushing stimuli similarly found that stroking was rated as pleasant even at 50 cm/s [36]. Furthermore, the apparent โU-shapedโ relationship between pleasantness and velocity may, in part, reflect averaging across individuals, as suggested by a recent review of five stroking paradigm studies in which only 40% of participants exhibited the canonical U-shaped pattern [26]. Thus, a focus on brush velocities in the โoptimal rangeโ of 1โ10 cm/s may represent the recipientโs perspective, whereas incorporating naturally observed patterns of deliverers may enhance the ecological realism of affective touch paradigms.
Notably, in the present study, the highest mean velocity recorded by an individual brusher was 20.33 cm/s, which was still rated as pleasant. This finding aligns with the idea that the range of 1 โ 10 cm/s represents a subset, rather than an exclusive, range of pleasantness. A study utilizing brushing stimuli similarly found that stroking was rated as pleasant even at 50 cm/s [36]. Furthermore, the apparent โU-shapedโ relationship between pleasantness and velocity may, in part, reflect averaging across individuals, as suggested by a recent review of five stroking paradigm studies in which only 40% of participants exhibited the canonical U-shaped pattern [26]. Thus, a focus on brush velocities in the โoptimal rangeโ of 1โ10 cm/s may represent the recipientโs perspective, whereas incorporating naturally observed patterns of deliverers may enhance the ecological realism of affective touch paradigms.
Observation of brushstroke forces confirm current practice
We found that participants applied force magnitudes of 0.37 ยฑ 0.24 N, consistent with typical ranges of 0.2 โ 0.4 N [14], [18], [19], [28], [30]. We note that we calculate force magnitude as an aggregate of forces along X, Y, and Z axes. In contrast, most prior studies have derived force from the normal force calibration of a rotary tactile simulator [14], [19], [25], [27], [32], i.e., a brush attached to a single motor to produce a sweeping motion. Such normal force corresponds to Z-axis force of the instrumented brush used herein. Since our experiment accounts for an aggregate of forces along the X, Y, and Z axes, it is reasonable that our reported force range is slightly higher, as it also includes lateral forces (Fig. 1C).
Moreover, the alignment of our force measurements and typical ranges may be attributed, in part, to the properties of the brush bristles. Similar to von Frey monofilaments, brush bristles act as a force-limiting mechanism, preventing excessive force application beyond a threshold [43]. A study examining the perception of brushing from various brushes, found that different brush types (goat hair, horse hair, plastic bristles) deliver varying force values, even when stroked in a similar manner [15]. This suggests brush stiffness naturally constrains force application, contributing to experimental consistency, offering a simple way to modulate force delivery.
Human Brusher Repeatability and Robotic Delivery
Within the delivery strategies of individuals, brushers demonstrated repeatable stroke delivery across trials โ as indicated by low coefficients of variation, non-significant results in Leveneโs tests, and high ICC values for both force and velocity โ without formal training. While the degree of precision varies between individuals, as observable in the relative variability in their cluster sizes, each participantโs variability was characteristic and internally consistent. For example, a highly precise brusher (Fig. 6: small cluster, Participant 4) and a less precise brusher (Fig. 6: large cluster, Participant 1) each could readily replicate their unique cluster sizes between their 5 trials.
Moreover, in attempt to regulate the velocity and force of human delivered stroking, prior studies have used metronomes [33], [34] or moving visual bars [29] and pre-training with force-calibrated scales [15]. Although we did not utilize or compare the precision of strokes with or without such aids, the results herein demonstrate that untrained participants are nevertheless consistent in their stroking delivery. Further, our study neither evaluated trained participants, nor their ability to deliver accurate force and/or velocity in accordance with a control condition. To address such questions, future studies might incorporate real-time measurement of stroking velocity and force to verify that commanded delivery parameters correspond with actual performance.
Beyond pre-study training and visual guidance to standardize touch delivery of human touch deliverers, rotary tactile simulators [13], [14], [19], [25], [27], [28], [32] and other robotic systems [31] have sought to enable repeatable touch delivery. As with human brushing, robotic systems are not immune to delivery errors, and their performance has rarely been directly compared with human-delivered stroking. While many rotary tactile stimulators record applied force and velocity, their sweeping actuation motion may introduce discrepancies between the commanded and actual dynamics and lack resemblance to human-delivered brushing. A recent study redesigned a robotic apparatus to incorporate a cord-driven mechanism to more closely replicate the natural trajectory of human motion [31]. This sort of systematic evaluation of robotic performance may reveal accuracy in reproducing target velocity and force profiles and approximating human stroking. In some cases, robotic touch may be overly consistent, or too โrobotic.โ The introduction of controlled variability may help robotic systems better emulate the natural dynamics of human touch in experimental contexts.
Gaussian Mixture Models for Touch Delivery Characterization
Within the delivery strategies of individuals, brushers demonstrated repeatable stroke delivery across trials โ as indicated by low coefficients of variation, non-significant results in Leveneโs tests, The lack of correlation between delivered force and velocity in pleasant brushing delivery was unexpected and, to our knowledge, has not been previously reported. Given their independence, we developed a two-dimensional clustering model to represent individualized brushing profiles within a force-velocity space (Figs. 5 and 6). Compared to traditional univariate methods, this approach offers a more nuanced framework for characterizing brushing strategies, capturing the complexity of individual behaviors across all five trials. Moreover, it provides interpretable visual outputs that facilitate direct comparisons between individuals.
To contextualize our approach, we drew inspiration from other multi-dimensional applications of GMMs in behavior classification, such as ambulatory monitoring, and individual path analysis for gesture recognition for telemanipulation [44], [45]. In ambulatory monitoring, GMMs have been used to classify behaviors such as falls using data from a single hip-worn accelerometer. For example, a 2006 study by Allen and colleagues could effectively classify postures and movements, e.g. sitting, standing, walking, standing-to-lying, etc. The study found that time-based features improved accuracy and reduced computational demand compared to frequency-based measures [44]. However, the resulting model relied on a 25-dimensional feature vector at each sample and a 32-mixture GMM, which limited direct interpretability of individual features. This โblack-boxโ nature constrains transparency and makes it difficult to understand why specific classifications are made. While such models can detect behavioral differences, they often do not explain the underlying mechanisms driving those differences. In contrast, our focus on both populationlevel and individual clustering while limiting complexity allows us to extract meaningful insights about the features themselves, including force and velocity ranges for the participant pool, as well as individual-specific ranges and consistency. For future touch-focused GMM applications, care should be taken to ensure that model decisions are interpretable, and that the contributions of individual features are clearly understood and contextualized in relation to both the population and individual participants.
Importantly, GMMs are a form of soft clustering, making them particularly well-suited for observation discovery for behaviors where multiple solutions are valid. A similar challenge arises in robot-assisted surgery, where complex procedures can be executed in multiple ways. Haptic feedback is often integrated into robotic systems to guide the surgeon along a predefined trajectory or provide interactive assistance. To better understand when and how this feedback should be deployed, Pรฉrez-del-Pulgar et al. applied GMMs to break down a multi-path task, inserting a peg into a hole, into stages, revealing the numerous valid strategies participants can use and identifying patterns that inform stepwise feedback [45]. The resulting GMM encoded surface contact patterns (e.g., upโdown, leftโright), enabling the system to detect consistent interactions across multiple trials and prioritize cues for haptic guidance. Observation-based experiments such as this and our study benefit from GMMsโ soft clustering, which allows patterns to emerge from complex, high-dimensional data without enforcing rigid classifications. In our touch delivery study, just as in robotic surgery, multiple brushing strategies may be equally effective, corresponding to individual clusters, while commonalities across participants reveal shared tendencies, reflected in our identified cluster range. By leveraging soft clustering, GMMs enable us to capture both individual variability and population-level agreement, revealing the structure of force and velocity patterns across participants. This approach allows meaningful patterns to be extracted from naturalistic behavior, providing insight into how untrained individuals deliver touch, all derived from direct observation rather than imposed experimental constraints.
๊ธฐ์กด ์ ์์ ์ด๊ฐ(affective touch) ์ฐ๊ตฌ์ ํ๊ณ
์์ ์ ์ค์ฌ์ ๋ฐ์ดํฐ ํธํฅ: ๊ธฐ์กด์ pleasant stroking paradigm(์๋ 1~10 cm/s, ํ ์ฝ 0.4 N)์ ํฐ์น๋ฅผ ๋ฐ๋ ์ฌ๋(์์ ์)์ ์๋ฆฌ์ /์ง๊ฐ์ ๋ฐ์(CT ์ ๊ฒฝ ์ฌ์ ํ์ฑํ ๋ฑ)์๋ง ์น์ฐ์ณ ๋์ถ๋ ์์น์ด๋ค.
์์ฐ์ค๋ฌ์ด ํ์์ ๋งฅ๋ฝ ์์ค: ์ ์ ํฐ์น๋ฅผ ์ ๊ณตํ๋ ์ฌ๋(์ก์ ์)์ด ์์ฐ์ค๋ฝ๊ฒ ํ๋ํ ๋ ์ด๋ค ๋ฌผ๋ฆฌ์ ์์น์ ์ ์ด ์ญํ์ผ๋ก ํฐ์น๋ฅผ ์ํํ๋์ง์ ๋ํ ์ค์ ๊ด์ฐฐ ๋ฐ์ดํฐ๊ฐ ์ฌ๊ฐํ๊ฒ ๋ถ์กฑํ๋ค.
์์์ ์ธ ์คํ ์กฐ๊ฑด์ ๊ณ ์ฐฉํ: ์ ํ ์ฐ๊ตฌ๋ค์์ ๋ค๋ฃจ์ด์ง ๊ณ ์ ์๋ ๊ธฐ์ค๋ค(์, 0.1, 0.3, 1, 3, 10, 30 cm/s)์ ์ด๊ธฐ ์คํ ์ค๊ณ ๋น์ ์ธ์์ ์ผ๋ก ์ค์ ๋ ๊ธฐ์ค์ผ ๋ฟ์ด๋ฉฐ, ์ด๊ฒ์ด ์ธ๊ฐ์ด ์์ฐ์ค๋ฝ๊ฒ ์ ํธํ์ฌ ์์ฐํ๋ ๋ฌผ๋ฆฌ์ ํจํด์ ์จ์ ํ ๋๋ณํ์ง ๋ชปํ๋ค.
๋จ์ผ ๋ณ์ ๋ถ์ ๋ฐ ๋ก๋ด ํต์ ์ ํ๊ณ: ํ๊ณผ ์๋๋ฅผ ๋
๋ฆฝ๋ ๋ณ๊ฐ์ ๋ณ์๋ก๋ง ์ทจ๊ธํ์ฌ ๊ฒฐํฉ๋ ๋ค์ฐจ์์ ํฐ์น ์ ๋ต์ ๋์ณค์ต๋๋ค. ๋ํ, ์คํ์ฉ ํ์ ์ด๊ฐ ์๊ทน๊ธฐ(๋ก๋ด)๋ ์ธ๊ฐ ๊ณ ์ ์ ์คํธ๋กํฌ ๊ถค์ ์ด๋ ์์ฐ์ค๋ฌ์ด ๊ฐ๋ณ์ฑ์ ๋ชจ์ฌ(emulating)ํ๋ ๋ฐ ํ๊ณ๊ฐ ์กด์ฌํ๋ค.
ํต์ฌ ๊ธฐ๋ฅ ๋ฐ ์ฐ๊ตฌ ๋ฐฉ๋ฒ๋ก
์ผ์๊ฐ ๋ด์ฅ๋ ์ ๋ฐ ๋ธ๋ฌ์(instrumented brush) ๊ฐ๋ฐ: 6์ถ ๋ก๋์ (ATI Nano17)๊ณผ ์ ์๊ธฐ ์์ด๋ก ํธ๋ํน ์์คํ (Flock of Birds)์ 3D ํ๋ฆฐํ ์์์ ๊ฒฐํฉํ์ฌ, ํฐ์น ์๊ฐ์ ํ, ํ ํฌ, 3์ฐจ์ ์์น, ์๋, ์คํธ๋กํฌ ๊ธธ์ด๋ฅผ 335 Hz์ ๊ณ ํด์๋๋ก ๋๊ธฐํ ๋ฐ ์ธก์ ํ๋ค.
์์ฐ์ค๋ฌ์ด ํฐ์น ๋ฐ์ดํฐ ์์ง(N=20): 20๋ช ์ ํ๋ จ๋ฐ์ง ์์ ์ฐธ๊ฐ์๊ฐ ์ปคํผ ๋ค์ ์จ๊ฒจ์ง ์์ ์์ ํ๋์ ์์ ์ด ์๊ฐํ๊ธฐ์ '๊ฐ์ฅ ๊ธฐ๋ถ ์ข์ ๋ฐฉ์(pleasant touch)'์ผ๋ก ๋ถ์ง์ ํ๋๋ก ์ ๋ํ์ต๋๋ค. ์ฐธ๊ฐ์๋ณ๋ก 5ํ ๋ฐ๋ณต ์๋(trials)๋ฅผ ๊ฑฐ์ณ ์ด 1,087๊ฐ์ ์ ํจ ์คํธ๋กํฌ(stroke) ํํ์ ์ถ์ถํ์ต๋๋ค.
2์ฐจ์ ๊ฐ์ฐ์์ ํผํฉ ๋ชจ๋ธ(GMM) ๊ธฐ๋ฐ ๋ถ์: ํฐ์น ์๋ ฅ(force)๊ณผ ์๋(velocity) ์ฌ์ด์ ์ ์๋ฏธํ ์ ํ ์๊ด๊ด๊ณ๊ฐ ์์์ ํต๊ณ์ ์ผ๋ก ํ์ธํ๊ณ , ์ด ๋ ๋
๋ฆฝ ๋ณ์๋ฅผ ๊ฒฐํฉํ์ฌ 2์ฐจ์ ๋ฐ๋ ๊ณต๊ฐ์์์ ๊ฐ์ธ๋ณ ํฐ์น ํ๋กํ์ ์ํํธ ํด๋ฌ์คํฐ๋ง(soft clustering) ๊ธฐ๋ฒ์ผ๋ก ์๊ฐํ ๋ฐ ๋ชจ๋ธ๋งํ์ต๋๋ค.
์ฐ๊ตฌ ๊ฒฐ๊ณผ
ํต๋ ๋ณด๋ค ๋น ๋ฅธ ์์ฐ์ค๋ฌ์ด ํฐ์น ์๋: ์ฐธ๊ฐ์๋ค์ ํ๊ท ํฐ์น ํ(0.37 N ยฑ 0.24)์ ๊ธฐ์กด ํ๊ณ์ ํ์ค(0.4 N)๊ณผ ์ ์ฌํ์ผ๋, ํ๊ท ์๋๋ 13.7 cm/s ยฑ 8.8๋ก ๊ธฐ์กด ์ต์ ๋ฒ์(1~10 cm/s)๋ณด๋ค ์ ์๋ฏธํ๊ฒ ๋๊ฒ ์ธก์ ๋๋ค. ์ฌ์ง์ด, ์ต๊ณ ์๋ ์์ญ์ธ 20.33 cm/s์กฐ์ฐจ ์์ ์๋ ์ฌ์ ํ '๊ธฐ๋ถ ์ข์'์ผ๋ก ํ๊ฐํ๋ค.
๊ฐ์ธ ๊ณ ์ ์ ํฐ์น ์ ๋ต ๋ฐ ๋์ ๋ฐ๋ณต์ฑ(repeatability): ํผํ์ ์ ์ฒด(cohort)์ ๋ฐ์ดํฐ ๋ถํฌ๋ ์ผ์ ํ ๋ฒ์ ์์ ๋ญ์ณ ์์์ผ๋, ๊ฐ์ธ๋ณ๋ก๋ ๊ทธ ์์์ ์์ ๋ง์ ๋งค์ฐ ์ข๊ณ ๊ณ ์ ํ ์์ญ(subsets)๋ง์ ์ ์ ํ๊ณ ์์๋ค. ๋ํ, 5ํ์ ๋ ๋ฆฝ๋ ์คํ ์ ๋ฐ์ ๊ฑธ์ณ ๋ณธ์ธ ๊ณ ์ ์ ํ/์๋ ์กฐํฉ๊ณผ ์ ๋ฐ๋(cluster size)๋ฅผ ์ ์งํ๋ ๋์ ์ผ๊ด์ฑ์ ์ฆ๋ช ํ๋ค. (๊ธ๋ด์๊ด๊ณ์ ICC ๊ฐ: ์๋ 0.986, ํ 0.964)
์ฐ๊ตฌ ํจ๋ฌ๋ค์ ์ ํ ์ ์: ์ธ๊ฐ์ด ์์ฐ์ค๋ฝ๊ฒ ํํ๋ ์ ์์ ํฐ์น์ ์ํํ์ ํ๋น์ฑ(ecological validity)์ ํ๋ณดํ๊ธฐ ์ํด, ํฅํ ์ด๊ฐ ์ฐ๊ตฌ๋ 10~30 cm/s ์ฌ์ด์ ์๋ ๊ตฌ๊ฐ์ ๋ ์ธ๋ฐํ๊ฒ ํ๊ฐํด์ผ ํ๋ฉฐ, ํ๊ณผ ์๋ ๋ฐ ๊ฐ๋ณ์ฑ์ ํตํฉ๋ ํ๋์ ๊ฐ์ธ ๋ธ๋ฌ์ฑ ์ ๋ต(brushing strategy)์ผ๋ก ๋ค๋ฃจ์ด์ผ ํจ์ ์์ฌํ๋ค.
์ฅ์
๊ณ ํด์๋ ๋ฉํฐ๋ชจ๋ฌ ์ฑํฌ: 6์ถ ํ/ํ ํฌ ์ผ์์ ์์น ์ถ์ ์ผ์๋ฅผ 335Hz ์ฃผ๊ธฐ๋ก ๋๊ธฐํํ์ฌ, ๋ฏธ์ธํ ์๋๋ฆผ ์์์ ์ค์๊ฐ ๋ฌผ๋ฆฌ ๊ฐ์ ๋ณํ๋ฅผ ์๊ณก ์์ด ์จ์ ํ ํฌ์ฐฉํ๋ค.
GMM ๊ธฐ๋ฐ ๋ค์ฐจ์ ํ๋กํ์ผ๋ง: ๋จ์ผ ๋ณ์ ํต๊ณ์ ํ๊ณ๋ฅผ ๋์ด, ๊ฐ์ฐ์์ ํผํฉ ๋ชจ๋ธ(GMM) ํด๋ฌ์คํฐ ๋ถ์์ ํตํด ํ์์์ ํฐ์น ์คํ์ผ(์, ํ์ ๊ฐํ์ง๋ง ๋๋ฆฌ๊ฒ, ํน์ ํ์ ์ฝํ์ง๋ง ๋น ๋ฅด๊ฒ ๋ฑ)์ ๋ค์ฐจ์ ๊ณต๊ฐ์์ ์๊ฐ์ /๊ฐ๊ด์ ์ผ๋ก ์ ๋ฐํ๊ฒ ์ ์ํ๋ค.
ํ๊ณ์
์ผ๋ฐฉํฅ์ ํ์ ๋ถ์: ๋ณธ ์คํ์ ํฐ์น๋ฅผ ์ ๊ณตํ๋ ์ฌ๋์ ๋ฌผ๋ฆฌ์ ํ๋์์ ๊ท๋ช ์ ์ง์ค๋์ด ์์ด, ์ก์ ์์ ์คํธ๋กํฌ ๋ณํ์ ๋ฐ๋ผ ํฐ์น๋ฅผ ๋ฐ๋ ์ฌ๋(receiver)์ ์ค์๊ฐ ๊ฐ์ ์ํ๋ ์๋ฆฌ์ ํผ๋๋ฐฑ์ด ์ด๋ป๊ฒ ๋์ ์ผ๋ก ๋ณํ๋์ง ์๋ฐฉํฅ(bidirectional) ๋งํฌ๋ฅผ ์๋ฒฝํ ์ค์๊ฐ ๋งคํํ์ง๋ ๋ชปํ๋ค.
๋๊ตฌ ๋ฐ ๋ถ์์ ์ ํ์ฑ: ์คํ์ด ๋ธ๋ฌ์๋ผ๋ ํน์ ๋งค๊ฐ์ฒด์ ์ธ์ฒด์ ํ๋์ด๋ผ๋ ์ ํ๋ ์ ์ฒด ๋ถ์๋ง์ ๋์์ผ๋ก ์งํ๋์๊ธฐ ๋๋ฌธ์, ์ค์ ์๊ฐ๋ฝ ์ง๋ฌธ์ ์ฐ๊ฑฐ๋ ๋ค๋ฅธ ์ ์ฒด ๋ถ์๋ฅผ ํฐ์นํ ๋์ ์ญํ ๊ด๊ณ๊น์ง 100% ์ผ๋ฐํํ๊ธฐ์๋ ๋ณ์๊ฐ ์กด์ฌํ๋ค.
ํฅํ๊ณผ์
ํผ๋ถ ๊ฐ(skin-to-skin) ๋ค์ด๋ด๋ฏน์ค ์ธก์ ๊ณ ๋ํ: ๋ธ๋ฌ์ ๋๊ตฌ๋ฅผ ๋์ด, ์ฐฉ์ฉํ ์ ์ฐ ์ผ์(flexible sensor skin)๋ ๋น์ ์ด์ ๊ณ ์ ๋ฐ ๊ดํ ์ถ์ ์ฅ์น๋ฅผ ๊ฒฐํฉํ์ฌ ๋๊ตฌ ์์ด ์ธ๊ฐ ๋ ์ธ๊ฐ์ ์์ ํผ๋ถ ๊ฐ ์ ์์ ํฐ์น ์ญํ์ ์์ ์ ํธ ๋ ๋ฒจ์์ ์ ๋ฐ ๋ถ์ํ ๊ณํ์ด๋ค.
์ฌํ์ /๊ด๊ณ์ ์ปจํ ์คํธ ๊ฒฐํฉ: ํฐ์น๋ฅผ ๋๋๋ ๋ ์ฌ๋ ์ฌ์ด์ ์ฌํ์ ๊ด๊ณ(๋ฏ์ ์ฌ๋, ์น๊ตฌ, ์ฐ์ธ, ๊ฐ์กฑ ๋ฑ)๋ ๋ํ์ ๋ฌธ๋งฅ(์๋ก, ์ถํ, ์น๋ฐ๊ฐ ํํ ๋ฑ)์ ๋ฐ๋ผ ๊ฐ์ธ์ ํ, ์๋, ํฐ์น ์ ๋ต๊ณผ ๋ฐ๋ณต์ฑ(repeatability)์ด ์ด๋ป๊ฒ ๊ฐ๋ณ์ ์ผ๋ก ๋ณ๋ชจํ๋์ง ์ฌ๋ฆฌ/๋ฌผ๋ฆฌ ํตํฉ ์ฐ๊ตฌ๋ฅผ ์งํํ ์์ ์ด๋ค.
Erzhen Hu, Qian Wan, Changkong Zhou, Md Aashikur Rahman Azim, PiaoHong Wang, Xingyi Hu, Yuhan Zeng, Zhicong Lu, Seongkook Heo
ThingMoji: User-Captured Cut-Outs For In-Stream Visual Communication
(Abstract) Live streaming has become increasingly popular, driven by the desire for direct and real-time interactions between streamers and viewers. However, current text-based interactions and pre-defined emojis limit expressiveness, especially when referring to specific stream moments. We propose ThingMoji, a type of user-captured cut-outs to enhance user expression and foster more effective communication between streamers and their audience in the comment section. ThingMojis are unique digital icons created by users by capturing snapshots and annotating specific areas at any point during the stream. We developed StreamThing, a live-streaming platform integrated with ThingMojis, to explore their use during object-focused live streaming contexts. In a user study with three in-the-wild deployments reveals the expressive use of ThingMojis in diverse live-streaming scenarios with rich visual contents. Our findings show that ThingMojis enable viewers to reference specific objects, express emotions, and create shared visual narratives. Streamers found ThingMojis valuable for facilitating on-the-fly communication around visual content and fostering playful interactions. The study also uncovered challenges in ThingMoji comprehension, issues for long-term uses of ThingMojis, and potential concerns regarding misuse. Based on these insights, we discussed new opportunities for supporting object-focused communication during live streaming environments.
(Introduction) Live streaming has become an integral part of online communication and is used across various domains such as gaming, crafting, and e-commerce [51]. What sets live streaming apart from other digital media is its real-time, dynamic nature, where streamers and audiences engage in a real-time exchange, fostering a sense of community and participation [4, 12, 19, 24, 25]. Viewers can comment and ask questions during the stream, allowing streamers gauge viewersโ interests and dynamically adjust the content and direction. Consequently, viewer-streamer communication has become a key aspect of the streaming experience, moving beyond passive video consumption.
A critical challenge in live streaming emerges from the inadequacy of current communication tools โ predominantly text chat and pre-defined emojis [4] โ in supporting precise object-focused discourse. While viewers frequently need to reference specific objects, moments, and spatial relationships within streams [17, 50, 62, 64, 72, 81], existing tools often fall short in facilitating these nuanced interactions. These communication challenges are especially evident in streams with complex visual content. For instance, when a jewelry maker demonstrates intricate wirewrapping techniques or when a model builder showcases various assembly components, viewers struggle to precisely indicate which elements theyโre discussing. Similarly, in e-commerce streams where multiple products are displayed simultaneously [72], viewers often resort to ambiguous temporal references (โthe necklace you showed earlierโ) or imprecise spatial descriptions (โthe blue piece on the leftโ). Such descriptions not only consume time to compose but frequently lead to miscommunication and reduced engagement between streamers and viewers. Such descriptions lack the details and can be time-consuming to write and difficult for other viewers and the streamer to quickly understand, resulting in communication gaps and reduced engagement [73].
Previous research has attempted to address these referencing challenges through various visual annotation approaches [49, 73]. These systems [49, 73] have integrated pre-defined and free-form annotation tools for users to attach snapshots and sketches with chat-based comments, aiming to provide a more direct and visual form of communication. However, these snapshot-based methods with screenshot attachment can clutter the comment section and disrupt focus. These methods often result in coarse-grained area of interest identification and disjointed interactions, as they lack integration with textual communication. Moreover, while these approaches have shown promise in creative live streams like visual arts [49, 73], they struggle with object-focused communication tasks [41, 50, 56], e.g., hands-on activities like maker projects, crafting, and model building, e-commerce interactions involving shared physical objects, and specific gaming streams that emphasize tracking objects of interest rather than general areas of interest.
We believe that addressing these challenges and facilitating fluid visual communication of objects in live streams is critical to ensuring effective communication between viewers and the streamer. To achieve this goal, our work draws inspiration from how digital icons and emojis have successfully enhanced expressivity in online communication [16, 63]. While emojis excel at conveying emotions and simple concepts succinctly, current live streaming platforms limit viewers to pre-defined sets of emojis that cannot capture or reference the dynamic visual content being shared [49]. This suggests an opportunity: combining the communicative efficiency of emojis with the ability to reference specific objects and moments from the stream.
Our work investigates the use of user-captured cut-outs, named ThingMojis, designed for object-focused communication. Unlike typical pre-defined emojis, ThingMojis allow users to efficiently communicate context in text messages while preserving the natural flow of conversation. ThingMoji offers a unique way to enhance live-stream interaction by representing key objects of interest. It is object-based, capturing items identified by viewers; context-based, preserving the whole imagery from the stream and allowing users to create narratives enriched with contextual meanings; and time-based, turning selected live-stream moments into expressive cut-outs aligned with the video timeline.
With design principles derived from prior literature, we developed a live streaming system, StreamThing, as a prototype to investigate the use of ThingMojis during object-focused live streaming contexts. StreamThing enables viewers to efficiently create ThingMojis by brushing over objects they want to reference in the live video. These ThingMojis can then be seamlessly embedded into chat messages like regular emojis, carrying with them the rich context of when and where they were captured in the stream. The system also helps viewers and streamers keep track of these visual references through a timeline that shows when different ThingMojis were used throughout the stream. We conducted a deployment study with three streamers and 29 viewers using StreamThing. Our investigation focused on understanding how viewers create and use ThingMojis, howthese visual elements integrate into stream communication, and how their meaning and utility evolve over time. The findings demonstrated that viewers used ThingMojis as an expressive visual language channel for referencing and self-expression, while also revealing important challenges in comprehension and organization that inform future design of similar systems.
In summary, our contributions include:
The design and implementation of ThingMoji in a live stream platform that streamlines referencing and communicating spatial and temporal contexts in live video.
Results from the deployment study about the use of ThingMoji for in-stream communication during object-focused live-streaming contexts.
(Conclusion) Weintroduce ThingMoji, a user-captured cut-out designed to enhance streamer-viewer communication. ThingMojis serve as a powerful tool for spotlighting, condensing, and summarizing crucial contextual information related to object-oriented elements across various application scenarios, including creative crafting, e-commerce product presentations, and gaming. Our user study with streamers and viewers revealed promising use cases for ThingMojis across diverse live-streaming contexts, with design implications for the continued exploration of the ThingMoji concept. We anticipate that ThingMoji will pave the way for novel opportunities in object-focused sharing within the realm of live streaming, sparking further exploration of such design principles within the HCI community.
๊ธฐ์กด ๋ผ์ด๋ธ ์คํธ๋ฆฌ๋ฐ ์ํต ๋ฐฉ์์ ํ๊ณ
๊ฐ์ฒด ์ค์ฌ ์ํต(object-focused discourse)์ ๋ถ์ฌ: ๋ณต์กํ ์๊ฐ ์ ๋ณด๊ฐ ํฌํจ๋ ๋ฐฉ์ก(์๊ณต์, ์ฅ๋๊ฐ ์กฐ๋ฆฝ, ์ด์ปค๋จธ์ค ๋ฑ)์์ ์์ฒญ์๊ฐ ํน์ ๋ถํ์ด๋ ์ ํ์ ์ง์นญํ ๋, ๊ธฐ์กด์ ๋จ์ ํ ์คํธ ์ฑํ ์ด๋ ๊ธฐ์ฑ ์ด๋ชจํฐ์ฝ ์ธํธ๋ก๋ ์ ํํ ๋ ผ์๊ฐ ๋ถ๊ฐ๋ฅํ๋ค.
๋ชจํธํ ์๊ณต๊ฐ์ ํํ์ ํ๊ณ: ์์ฒญ์๋ค์ "์๊น ์ผ์ชฝ์ ๋ณด์ฌ์ค ํ๋์ ๋ชฉ๊ฑธ์ด์"์ ๊ฐ์ด ๋ชจํธํ๊ณ ๊ธด ์ค๋ช ์ ์์ฑํด์ผ ํ๋ฏ๋ก ์์ฑ ์๋๊ฐ ํ์ด๋ฐ์ ๋์น๊ธฐ ์ฝ๊ณ , ์คํธ๋ฆฌ๋จธ์ ๋ค๋ฅธ ์์ฒญ์๊ฐ ์ด๋ฅผ ์ฆ๊ฐ์ ์ผ๋ก ์ดํดํ๊ธฐ ์ด๋ ค์ ์ค์๊ฐ ์ํต์ ๋จ์ ์ด ๋ฐ์ํ๋ค.
๊ธฐ์กด ์๊ฐ์ ์ฃผ์(annotation) ๋ฐฉ์์ ๋ฌธ์ ์ : ํ๋ฉด ์ ์ฒด ์คํฌ๋ฆฐ์ท์ ์ฒจ๋ถํ๊ฑฐ๋ ํ๋ฉด ์์ ๊ทธ๋ฆผ์ ๊ทธ๋ฆฌ๋ ๋ฐฉ์์ ์ฑํ
์ฐฝ์ ๊ณผ๋ํ๊ฒ ๊ฐ๋ ค ์๊ฐ์ ๊ณตํด(visual clutter)๋ฅผ ์ ๋ฐํ๊ณ , ๋ํ์ ์์ฐ์ค๋ฌ์ด ํ๋ฆ์ ๋ฐฉํดํ๋ฉฐ, ํ
์คํธ ๋ฉ์์ง์์ ์ ๊ธฐ์ ์ธ ์ธ๋ผ์ธ(inline) ๊ฒฐํฉ์ด ๋ถ๊ฐ๋ฅํ๋ค.
ThingMoji(StreamThing ์์คํ )
๋ผ์ด๋ธ ์คํธ๋ฆฌ๋ฐ ์ค ํ ์คํธ ์ฑํ ๊ณผ ๊ณ ์ ๋ ์ด๋ชจํฐ์ฝ์ ํ๊ณ๋ฅผ ๊ทน๋ณตํ๊ธฐ ์ํด, ์์ฒญ์๊ฐ ์ค์๊ฐ ๋ฐฉ์ก ํ๋ฉด ์ ํน์ ์ค๋ธ์ ํธ๋ฅผ ๋ฌธ์ง๋ฌ ๋๋ผ(cut-out)๋ฅผ ๋ฐ๊ณ ์ด๋ฅผ ์ด๋ชจํฐ์ฝ์ฒ๋ผ ์ฑํ ์ฐฝ์ ์ฆ๊ฐ ์ฝ์ ํ์ฌ ์ํตํ ์ ์๊ฒ ํ ํ์ ์ ์ธ ๋ผ์ด๋ธ ์คํธ๋ฆฌ๋ฐ ํ๋ซํผ์ด๋ค.
์ฌ์ฉ์ ์ ์ํ ์ปท์์ ThingMoji: ์์ฒญ์๊ฐ ์ค์๊ฐ ๋ฐฉ์ก ์ค ์ํ๋ ๊ฐ์ฒด ์๋ฅผ ๋ฌธ์ง๋ฅด๋ฉด(marking/brushing), ๋ชจ๋ฐ์ผ ๊ฒฝ๋ ์ธ๊ทธ๋ฉํ ์ด์ ๋ชจ๋ธ(MobileSAM) ์๊ณ ๋ฆฌ์ฆ์ด ๊ฐ๋๋์ด ํด๋น ์ค๋ธ์ ํธ๋ง ํฌ๋ช ํ ์์ด์ฝ ํํ๋ก ์ค์๊ฐ ์ถ์ถ(cut-out)ํ๋ค.
๊ฐ์ฒด/๋งฅ๋ฝ/์๊ฐ ๊ธฐ๋ฐ ๋ฐ์ดํฐ ๋ฐ์ธ๋ฉ: ThingMoji๋ ๋ ๋ฆฝ๋ ์ด๋ฏธ์ง๊ฐ ์๋๋ผ, ํด๋น ๊ฐ์ฒด๊ฐ ๋ฐฉ์ก ํ๋ฉด์ ์ด๋ ์์น(spatial context)์, ๊ทธ๋ฆฌ๊ณ ์ด๋ค ๋น๋์ค ํ์์คํฌํ(temporal context)์ ๋ฑ์ฅํ๋์ง์ ๋ํ ๋ฉํ๋ฐ์ดํฐ์ ์๋ณธ ์ค๋ ์ท ํ๋ ์ ์ ๋ณด๋ฅผ ๋ดํฌํ๋ค.
ํ
์คํธ-์๊ฐ ๊ฒฐํฉํ ์ธ๋ผ์ธ ์ฑํ
: ์์ฑ๋ ThingMoji ์์ด์ฝ์ ์ผ๋ฐ ๋ด์ฅ ์ด๋ชจํฐ์ฝ์ฒ๋ผ ํ
์คํธ ๋ฉ์์ง ์ฌ์ด์ ์์ฐ์ค๋ฝ๊ฒ ์๋ฒ ๋ฉํ์ฌ ์ ์กํ ์ ์๋ค. ์คํธ๋ฆฌ๋จธ๋ ๋ค๋ฅธ ์์ฒญ์๊ฐ ์ฑํ
์ฐฝ์ ThingMoji๋ฅผ ํด๋ฆญํ๋ฉด, ํด๋น ๊ฐ์ฒด๊ฐ ์ถ์ถ๋ ์๋ ๋ฐฉ์ก ํ๋ฉด๊ณผ ํ์๋ผ์ธ ์์ ์ด ํ์
๋์ด ๋งฅ๋ฝ์ ์ฆ๊ฐ ์ญ์ถ์ (context retrieval)ํ ์ ์๋ค.
์ฐ๊ตฌ ๊ฒฐ๊ณผ ๋ฐ ๋ฐฐํฌ ๊ฒ์ฆ
์ผ์ธ ์ค์ ๋ฐฐํฌ ์ฐ๊ตฌ(in-the-wild study, N=32): ์์คํ ์ ํ๋น์ฑ ๊ฒ์ฆ์ ์ํด 3๋ช ์ ์คํธ๋ฆฌ๋จธ(ํ๋ผ๋ชจ๋ธ ์กฐ๋ฆฝ, ๋ณด๋๊ฒ์, ๊ฐ๊ตฌ ๊ฐ๊ณต ๋๋ฉ์ธ)์ 29๋ช ์ ์ค์ ์์ฒญ์๋ฅผ ๋์์ผ๋ก ์ค์ ๋ผ์ด๋ธ ์คํธ๋ฆฌ๋ฐ ๋ฐฐํฌ ์คํ์ ์งํํ๋ค.
์๊ฐ์ ์ฐธ์กฐ ๋ฐ ๋ด๋ฌํฐ๋ธ ํ์ฑ: ์คํ ๊ธฐ๊ฐ ๋์ ์ด 193๊ฐ์ ๊ณ ์ ํ ThingMoji๊ฐ ์์ฑ๋์๊ณ ์ฑํ ์ฐฝ์์ ์ด 452ํ ์ฌ์ฉ๋์์ต๋๋ค. ์์ฒญ์๋ค์ ๋ณต์กํ ์ค๋ช ์์ด๋ ํน์ ์ค๋ธ์ ํธ๋ฅผ ์ฆ๊ฐ referencingํ ์ ์์์ผ๋ฉฐ, ์ด๋ฅผ ํ์ฉํด ์ ์พํ meme์ ๋ง๋ค๊ฑฐ๋ ๊ณต์ ๋ ์๊ฐ์ ๋ด๋ฌํฐ๋ธ(shared visual narratives)๋ฅผ ํ์ฑํ๋ค.
์คํธ๋ฆฌ๋จธ์ ์ฆ๊ฐ์ ์ํต ๋ฐ ์ ํฌ์ฑ ๊ฐํ: ์คํธ๋ฆฌ๋จธ๋ค์ ํ๋ฉด ์ ์ฝํ
์ธ ๋ฅผ ๊ธฐ๋ฐ์ผ๋ก ์์ฒญ์๋ค๊ณผ ์ฆ๊ฐ์ (on-the-fly)์ผ๋ก ์ํตํ๊ณ ํผ๋๋ฐฑ์ ์ค ์ ์๋ค๋ ์ ์์ ๋์ ๋ง์กฑ๋๋ฅผ ๋ณด์์ผ๋ฉฐ, ThingMoji๋ฅผ ํ์ฉํ ์ฅ๋์ค๋ฝ๊ณ ์ ์พํ ์ธํฐ๋์
(playful interactions)์ด ๋ฐฉ์ก ๋ชฐ์
๋๋ฅผ ๊ทน๋ํํจ์ ํ์ธํ๋ค.
์ฅ์
์ปค๋ฎค๋์ผ์ด์ ์ค๋ฒํค๋ ๊ฐ์: visual referencing๊ฐ ์ค์๊ฐ์ผ๋ก ์ด๋ฃจ์ด์ง๋ฏ๋ก, ํ์ดํ ์๋๊ฐ ๋๋ฆฌ๊ฑฐ๋ ๋ชจ๋ฐ์ผ ํ๊ฒฝ์ธ ์์ฒญ์๋ ๊ธด ์์ ํ ๋ฌธ์ฅ ์์ฑ ์์ด ์๋๋ฅผ ๋ช ํํ๊ฒ ์ ๋ฌํ ์ ์๋ค.
ํ๋ฉด ๊ณต๊ฐ์ ํจ์จ์ ํ์ฉ: ์ ์ฒด ์คํฌ๋ฆฐ์ท์ ๊ณต์ ํ๋ ๋ฐฉ์๊ณผ ๋ฌ๋ฆฌ, ๊ฐ์ฒด๋ง ์๋ผ๋ธ ์ปดํฉํธํ ์ธ๋ผ์ธ ์์ด์ฝ ํํ๋ฅผ ์ทจํ๋ฏ๋ก ํ์ ๋ ๋ชจ๋ฐ์ผ ์ฑํ ์ฐฝ ๋ ์ด์์์ ์๊ฐ์ ๊ณตํด(clutter)๋ฅผ ์ต์ํํฉ๋๋ค.
๋งฅ๋ฝ ๋ณด์กดํ ์ธํฐ๋์
: ๋จ์ํ ์ด๋ฏธ์ง ์คํฐ์ปค๋ฅผ ๋์ด, ํด๋ฆญ ์ ์๋ณธ ํ๋ ์์ ์๊ณต๊ฐ์ ์์น๋ฅผ ๋ณต์ํด ์ฃผ๋ฏ๋ก ๋ฐฉ์ก ์ค๊ฐ์ ์๋ก ๋ค์ด์จ ์์ฒญ์๋ ๋ํ ๋ฌธ๋งฅ์ ์ฝ๊ฒ ์ฐธ์ฌํ ์ ์๋ค.
ํ๊ณ์
์๋ฏธ์ ๋ชจํธ์ฑ(Comprehension Issue): ์ฃผ๋ณ ํ ์คํธ ๋งฅ๋ฝ์ด๋ ์ค๋ช ์ด ๋ถ์กฑํ ์ํ์์ ThingMoji ์์ด์ฝ๋ง ๋จ๋ ์ผ๋ก ์ฐ๋ฌ์ ์ ์ก๋ ๊ฒฝ์ฐ, ์คํธ๋ฆฌ๋จธ๋ ๋ค๋ฅธ ์์ฒญ์๊ฐ ํด๋น ์์ด์ฝ์ ์ ๋ณด๋๋์ง ์๋(์ธ์ฆ, ์ง๋ฌธ, ๋จ์ ๊ฐํ ๋ฑ)๋ฅผ ์ง๊ด์ ์ผ๋ก ํ์ ํ๊ธฐ ์ด๋ ค์ด ๊ฒฝ์ฐ๊ฐ ๋ฐ์ํฉ๋๋ค.
์์ด์ฝ ๊ด๋ฆฌ ๋ฐ ์กฐ์งํ ๋ถ๋ด: ํ ๋ฐฉ์ก์์ ์์ญ ๊ฐ์ ThingMoji๊ฐ ํ๊บผ๋ฒ์ ์์ฑ๋๋ฉด ์์ฒญ์์ '๋ด ThingMoji ๋ณด๊ดํจ' ํญ์ด ๋ณต์กํด์ ธ, ์ํ๋ ์ปท์์์ ๋น ๋ฅด๊ฒ ์ฐพ์ ์ฌ์ฌ์ฉํ๊ธฐ ์ด๋ ค์์ง๋ ๊ด๋ฆฌ ํจ์จ์ฑ ๋ฌธ์ ๊ฐ ๋ํ๋ฉ๋๋ค.
์
์ฉ(misuse) ๋ฐ ๊ฒ์ด ์์คํ
๋ถ์ฌ: ์คํธ๋ฆฌ๋จธ๊ฐ ์จ๊ธฐ๊ณ ์ถ์ด ํ๊ฑฐ๋ ๋ฐฉ์ก ์ฌ๊ณ ๋ก ๋
ธ์ถ๋ ๋ถ์ ์ ํ ํ๋ฉด ์์ญ, ํน์ ํ์ธ์ ํ๋ผ์ด๋ฒ์๊ฐ ์นจํด๋ ์ ์๋ ์ค๋ธ์ ํธ๋ฅผ ์์ฒญ์๊ฐ ์
์์ ์ผ๋ก ๋๋ผ ๋ฐ์ ๋๋ฐฐํ ๊ฒฝ์ฐ ์ด๋ฅผ ์ค์๊ฐ์ผ๋ก ์ ์ดํ ํํฐ๋ง ์ฅ์น๊ฐ ๋ถ์กฑํ๋ค.
ํฅํ๊ณผ์
์ค๋งํธ ์ ์ ๋ฐ ์๋ ์กฐ์งํ ์๊ณ ๋ฆฌ์ฆ: ์์ฒญ์๊ฐ ์ผ์ผ์ด ๋ฌธ์ง๋ฅด์ง ์์๋ ์คํธ๋ฆฌ๋จธ๊ฐ ํ์ฌ ๋ง์ง๊ณ ์๊ฑฐ๋ ์นด๋ฉ๋ผ ์ค์ฌ์ ์๋ ์ฃผ์ ๊ฐ์ฒด๋ฅผ AI๊ฐ ํ๋จํด ์๋จ์ ์ปท์์ ํ ํ๋ฆฟ์ผ๋ก ์๋ ์ ์(proactive recommendations)ํ๊ณ , ์นดํ ๊ณ ๋ฆฌ๋ณ๋ก ์๋ ๋ถ๋ฅํด ์ฃผ๋ ๋ณด๊ดํจ ๊ธฐ๋ฅ์ ๊ณ ๋ํํ ๊ณํ์ด๋ค.
๋ค์ค ํ๋ซํผ ์ค์ผ์ผ๋ง ๋ฐ ๋ณด์ ๊ฒ์ด ์ตํฉ: ์์ฒ ๋ช ์ด์์ด ์ฐธ์ฌํ๋ ๋๊ท๋ชจ ํธ์์น, ์ ํ๋ธ ๋ผ์ด๋ธ ํ๊ฒฝ์์๋ ์ฑํ ์ฑ๋ฅ ์ ํ(latency)๊ฐ ์๋๋ก ์ค์ผ์คํธ๋ ์ด์ ์ ์ต์ ํํ๊ณ , ์ปดํจํฐ ๋น์ ๊ธฐ๋ฐ์ ์ค์๊ฐ ์ ํด ๊ฐ์ฒด ๊ฒ์ด(NSFW/privacy filtering) ๋ชจ๋์ ๊ฒฐํฉํ์ฌ ์์ ํ ์๊ฐ ์ํต ์ธํฐํ์ด์ค๋ฅผ ์ ๋ฆฝํ ์์ ์ด๋ค.
Ajwa Shahid, Jane Chung, Seongkook Heo
Exploring Older Adults Personality Preferences for LLM-powered Conversational Companions
(Abstract) Studies have shown promise in using conversational agents (CAs) to reduce loneliness among older adults. Recent advances in large language models (LLMs) have enhanced these systems by enabling more natural, human-like interactions. However, little is known about how CA personalities influence user experiences, despite personality being a key factor in human conversation. To explore this, we developed a smart speaker-based CA powered by an LLM and conducted a two-phase user study consisting of in-lab sessions and home deployments. We investigated how older adults perceive different CA personalities and how these personalities affect their interaction experiences. Our findings show that participants could distinguish between different personality characteristics and had varying preferences for different personalities during both shortterm and long-term interactions.
(Introduction) Studies have shown that older adults are generally more susceptible to loneliness than other age groups, due to loss of social connections, declining health and mobility that limit social activities, and reduced interactions with family and friends [13]. A study also showed that loneliness has been linked with poor health and wellness outcomes for older adults [4]. As a result, there is growing interest within the research community in developing interven-tions to address loneliness in this population. One promising area of research in Human-Computer Interaction focuses on conversational agents (CAs) to alleviate loneliness among older adults. Research has shown that such CAs can reduce feelings of loneliness by providing easily accessible social interaction, especially for older adults who are more likely to have limited opportunities for social engagement [8]. Multiple studies have shown some promise in the potential of companion CAs in reducing loneliness among older adults [2, 20]. As this is an emerging field, most prior studies have focused on exploring the use of various CAs by older adults, including voice assistants like Amazon Alexa [33] and chatbots such as ChatGPT[1]. Limited research has explored how personalizing CAs can enhance user experience, with most existing studies focusing on contextualizing agent responses based on user information like preferences or routines [2]. However, many important aspects of personalization remain underexplored. One such aspect is the CAโs personality, as psychological studies show that personality traits play a significant role in shaping communication patterns [25].
Building on these insights, this work aims to explore the effects of different personalities of Large language model (LLM)-powered CAs on the user experience and perceptions of older adults to answer the following research questions:
RQ1: How do older adults perceive different personality traits in LLM-powered voice conversational agents?
RQ2: What are the observed effects of different personality traits in LLM-powered voice conversational agents on older adultsโ experience with using a conversational agent?
To answer these research questions in a naturalistic setting, we implemented a smart speaker device that includes a mini PC, speaker, and microphone sothat users can useit intheir homes, similar to the setup used in a previous study on the use of CAs [1]. We also developed a CA program that uses an LLM for processing the conversations. The CA system we developed supports three distinct personality profiles, each characterized by a dominant trait: Extroversion, Agreeableness, or Conscientiousness. To understand how CAโs personality traits affect older adultsโ perceptions and experiences, we conducted a two-phase user study. In the first phase, older adults participated in an in-lab session, engaging in a 10-minute conversation with each personality variant of the agent. The second phase involved a deployment study, where the device was set up in participantsโ homes for a 12-day period. Throughout the study, we examined participantsโ perceptions and experiences with each personality type using a multi-method approach, including personality assessments, preference ratings, interviews, daily diary entries, and analysis of conversation transcripts between the older adults and the agents. We found that most participants preferred an Agreeable CAin shorter conversations while preferring an Extroverted CA for long-term use. We also found that Agreeableness can sometimes be perceived as sycophantic, while Extroversion may come across as insincerely positive. We discuss how these traits can be balanced to make them more suitable for conversational companions. We then offer insights into potential next steps that can help achieve our goal of designing the ideal conversational companions for older adults.
(Conclusion) We observed that personality perceptions and preferences varied significantly among participants, with interesting differences observed across the contexts of the Lab Study and the Deployment Study. For example, participants who favored a particular personality during the short conversations in the Lab Study sometimes expressed a preference for a different personality during the longerterm Deployment Study. Phase 1 and 2 had 5 and 3 participants respectively, and while this is not a large enough sample size to understand participant experiences in depth, the findings from this study provide insights into how different personalities can be balanced to better suit the role of conversational companions for older adults across different contexts.
The next natural step in our research is to conduct a more thorough study with more participants and a longer deployment. These future investigations could also explore how the personalities of older adult participants influence their interaction experiences with CAsexhibiting distinct personalities, as previous studies have found relationships between participant characteristics and their perceptions of CAs [6, 35, 36, 38]. The findings from in-depth investigations will inform the design of a conversational agent as a conversation companion for older adults. This will allow us to assess how effectively these interactions help older adults reduce loneliness. Another promising direction could be identifying and tailoring a personality for a conversational agent, and then assessing the effectiveness of the CAโs interactions with the older adult participants in mitigating loneliness. This could be achieved by conducting pre and post-study loneliness assessments for over an extended deployment period.
Additionally, drawing from prior research and our observations that older adults value responses contextualized by previous interactions, we can consider developing a dynamic system capable of adapting its personality based on retained context. This approach would reduce the possibility of older adults having to frequently adjust personality settings, minimizing cognitive load to ensure a smoother user experience. Such a system has the potential to enhance both the usability and the emotional connection between older adults and their conversational agents.
Our eventual goal is to design personalized CAs for older adults that can alleviate loneliness and encourage or facilitate socialization with other people. These CAs should be seen as complementary tools rather than substitutes for human companionship. In this paper we take the first step toward that goal, presenting preliminary work from a multi-phase qualitative study on evaluating different induced personalities in LLM-powered CAs for older adults. We present initial findings on how older adults perceived and interacted with the LLM-powered CAs, each featuring a distinct induced personality. Based on our insights, we give recommendations on the design of personalized companion CAs for older adults and outline possible next steps to achieve our goal.
๊ธฐ์กด ๋ํํ ์์ด์ ํธ(CA) ์ฐ๊ตฌ์ ํ๊ณ
๊ฐ์ธํ ์์์ ๋ค์ฐจ์์ ํ์ ๋ถ์กฑ: ๊ธฐ์กด ๊ณ ๋ น์ธต ๋์์ CA ๊ฐ์ธํ ์ฐ๊ตฌ๋ ์ฃผ๋ก, ์ฌ์ฉ์์ ์ผ์ ์ด๋ ์ทจํฅ ๋ฑ ์ปจํ ์คํธ(๋งฅ๋ฝ) ์ ๋ณด๋ฅผ ๋จ์ ๋ฐ์ํ๋ ๋ฐ ๊ทธ์ณค์ผ๋ฉฐ, ์ธ๊ฐ ๋ํ์์ ํต์ฌ ์ญํ ์ ํ๋ ์์ด์ ํธ ์์ฒด์ personality๊ฐ ๋ฏธ์น๋ ์ํฅ์ ๊ฑฐ์ ์ฐ๊ตฌ๋์ง ์์๋ค.
๊ธฐ์กด ์์ฑ๋น์์ ๋งฅ๋ฝ์ ์ง ํ๊ณ: ์๋ง์กด ์๋ ์ฌ(Alexa) ๋ฑ ๊ธฐ์กด์ ์์ฉ ์์ฑ ๋น์๋ ์ด์ ๋ํ์ ๋ฌธ๋งฅ์ ๊ธฐ์ตํ์ง ๋ชปํด ์์ฐ์ค๋ฌ์ด ํ์ ์ง๋ฌธ(follow-up questions)์ ๋์ง์ง ๋ชปํ๋ฏ๋ก, ๊ณ ๋ น์ธต ์ฌ์ฉ์๋ค์ด ๋ํ์ ๋จ์ ๊ณผ ์ฌ๊ฐํ ์ข์ ๊ฐ์ ๋๋ผ๋ ์์ธ์ด ๋์๋ค.
์ฅ/๋จ๊ธฐ์ ์ฑ๊ฒฉ ํจ๊ณผ ๊ฒ์ฆ์ ๋ถ์ฌ: LLM์ ํ์ฉํด ๋ค์ํ ์ฑ๊ฒฉ ํน์ฑ์ ์ ๋(induction)ํ ์ ์์์ ์ฆ๋ช
๋์์ผ๋, ์ด๋ฌํ ์ธ์์ ์ธ ์ฑ๊ฒฉ ๋ณํ์ด ์ค์ ์ธ๋ก์์ ๊ฒช๋ ๊ณ ๋ น์ธต์ ์ฅ๊ธฐ์ ์ธ ์ค์ํ ํ๊ฒฝ(in-home deployment)์์ ์ด๋ค ๊ตฌ์ฒด์ ์ธ ์ฃผ๊ด์ ๊ฒฝํ ๋ณํ๋ฅผ ์ฃผ๋์ง ๊ฒ์ฆ๋ ๋ฐ์ดํฐ๊ฐ ์์๋ค.
ํต์ฌ ๊ธฐ๋ฅ ๋ฐ ์ฐ๊ตฌ ๋ฐฉ๋ฒ๋ก
LLM ๊ธฐ๋ฐ ์ฝํ ์คํธ ๋ฉ๋ชจ๋ฆฌ ์ค๋งํธ ์คํผ์ปค ๊ฐ๋ฐ: Raspberry Pi(๋๋ PC ๋ฒ ์ด์ค ์์คํ ), ์คํผ์ปค, ๋ง์ดํฌ, ์ ์ฉ ๋ฒํผ ๋ฐ ๋คํธ์ํฌ ๋ชจ๋์ 3D ํ๋ฆฐํ ์์์ ํ์ฌํ์ฌ ์ ์ํ๋ค. GPT-4o๋ฅผ ์ฐ๋ํ์ฌ ์ธ์ ์ด ๋๋ ๋๋ง๋ค ๋ํ ๋ด์ฉ์ ๊ตฌ์กฐํ๋ ์์ฝ๋ณธ(summary)์ผ๋ก ์๋ฒ ๋ฉ ์ ์ฅํด ๋ค์ ๋ํ ์ ์ด์ ๊ธฐ์ต์ ์๋ฒฝํ ๋ณต์(memory retention)ํ๋ ํ๋์จ์ด-์ํํธ์จ์ด ํตํฉ ์์คํ ์ ๊ตฌ์ถํ๋ค.
Pยฒ (Personality Prompting)๋ฅผ ํตํ ์ฑ๊ฒฉ ์ ๋: ๋น ํ์ด๋ธ(big five) ์ฑ๊ฒฉ ๋ชจ๋ธ ์ค ๋ถ์ ์ ๊ฐ์ ์ ์ ๋ฐํ๋ ์ ๊ฒฝ์ฆ(N)๊ณผ, ์ธํฅ์ฑ๊ณผ ๊ฐ๋ ์ด ์ค๋ณต๋๊ธฐ ์ฌ์ด ๊ฐ๋ฐฉ์ฑ(O)์ ์ ์ธํ๊ณ , ์ธํฅ์ฑ(E, extroverted), ์นํ์ฑ(A, agreeable), ์ฑ์ค์ฑ(C. conscientious) ์ธ ๊ฐ์ง ํ๋กํ์ ํ๋กฌํํธ ์์ง๋์ด๋ง ๊ธฐ๋ฒ์ผ๋ก ์ ๋ฐ ์ ๋ํ๋ค.
์คํ์ค(lab) ๋ฐ ๊ฐ์ ๋ด ๋ฐฐํฌ(deployment) 2๋จ๊ณ ์ฐ๊ตฌ: 1๋จ๊ณ๋ก 5๋ช
์ ๊ณ ๋ น์ธต ์ฐธ๊ฐ์์ ์คํ์ค์์ ๊ฐ ์ฑ๊ฒฉ๋ณ ์์ด์ ํธ์ 10๋ถ๊ฐ ๋ํํ๊ฒ ํ ํ(๋ผํด ์คํ์ด ๊ต์ฐจ ๊ฒ์ฆ), 2๋จ๊ณ๋ก 3๋ช
์ ์ฐธ๊ฐ์ ๊ฐ์ ์ ๊ธฐ๊ธฐ๋ฅผ ์ง์ ์ค์นํด 12์ผ๊ฐ ์ค์ ์ถ ์์์ ์ฅ๊ธฐ์ ์ธ ์ํธ์์ฉ ๋ฐ ์ผ๊ธฐ ๊ธฐ๋ก, ์ธํฐ๋ทฐ ๋ฐ์ดํฐ๋ฅผ ์์งํ๋ ํผํฉ ์ฐ๊ตฌ ๋ฐฉ๋ฒ๋ก ์ ์ฑํํ๋ค.
์ฐ๊ตฌ ๊ฒฐ๊ณผ
์ํธ์์ฉ ๊ธฐ๊ฐ์ ๋ฐ๋ฅธ ์ฑ๊ฒฉ ์ ํธ๋์ ๋ฐ์ : ๋จ๊ธฐ ๋ํ(์คํ์ค ์ธ์ )์์๋ ๋๋ถ๋ถ์ ์ฐธ๊ฐ์๊ฐ ๋ฐ๋ปํ ์นญ์ฐฌ๊ณผ ๊ณต๊ฐ์ ๊ฑด๋ค๋ ์นํ์ ์ธ(agreeable) ์์ด์ ํธ๋ฅผ ๊ฐ์ฅ ์ ํธํ๋ค. ํ์ง๋ง 12์ผ๊ฐ์ ์ฅ๊ธฐ ๋ํ(๊ฐ์ ๋ฐฐํฌ ์ธ์ )๋ก ๋์ด๊ฐ๋ฉด์, ๋จ์ ๋ฆฌ์ก์ ์ ๋์ด ์ฐฝ์์ ์ด๊ณ ์ ๊ทน์ ์ผ๋ก ๋ํ๋ฅผ ์ฃผ๋ํ๋ ์ธํฅ์ ์ธ(extroverted) ์์ด์ ํธ์ ๋ํ ์ ํธ๋๊ฐ ์ ์๋ฏธํ๊ฒ ๋์์ง๋ ์ญ๋์ ๋ณํ๋ฅผ ํ์ธํ๋ค.
์ฑ๊ฒฉ๋ณ ํ์ธ ํฌ์ธํธ(pain point) ๋ฐ๊ตด
์นํํ ์์ด์ ํธ: ์ฌ์ฉ์์ ๋ชจ๋ ๋ง์ ๋ฌด์กฐ๊ฑด ๋์กฐํ๊ณ ๋ํ์ดํ๋ ์์ฒจ(sycophancy) ๋ฐ ๋ฏธ๋ฌ๋ง ๊ฒฝํฅ์ด ๋์ ๋์ด, ์ฅ๊ธฐ์ ์ผ๋ก๋ "์ง์ง ์ธ๊ฐ๋ค์ด ๋ํ๊ฐ ์๋๋ผ ๋ก๋ด ๊ฐ๋ค"๋ผ๋ ์ง๋ฆผ ํ์์ ์ ๋ฐํ๋ค.
์ธํฅํ ์์ด์ ํธ: ๋์ ํ๋ ฅ๊ณผ ์ฆํฅ์ฑ(spontaneity)์ผ๋ก ์์์ ์์๋์ผ๋, ์ง๋์น ๋๊ด์ฃผ์์ ๊ณผ์ฅ์ผ๋ก ์ธํด ๋๋ฆฌ์ด ์ง์ ์ฑ ๊ฒฐ์ฌ(insincerity)๋ก ๋ค๊ฐ์ ์ฅ๊ธฐ์ ์ธ ์ ๋ขฐ ๊ด๊ณ๋ฅผ ์ ํดํ ์ ์์์ด ํ์ธ๋์๋ค.
์ฑ์คํ ์์ด์ ํธ: 3๊ฐ ์ฑ๊ฒฉ ์ค ๊ฐ์ฅ ๋ฎ์ ์ ํธ๋๋ฅผ ๊ธฐ๋กํ์ผ๋ฉฐ, ์ ๋ณด ์ ๋ฌ ์์ฃผ์ ๋ฑ๋ฑํ๊ณ ์ ํํ๋ ๋ํ ์คํ์ผ๋ก ์ธํด ๊ณ ๋ น์ธต์๊ฒ ๋ฌด๊ด์ฌํ๊ณ ์ง๋ฃจํ(disengaged/blah) ์ฑ๊ฒฉ์ผ๋ก ์ธ์๋์๋ค.
์์ฑ ๋ฐ ์ฑ๋ณ ํ ๋น์ ๊ต์ฐจ ํจ๊ณผ(gender assignment): ์์คํ
์์ ์ฑ๋ณ ์ค๋ฆฝ์ (gender-neutral) ์์ฑ์ ์ ๊ณตํ์์๋, ์ฐธ๊ฐ์๋ค์ ์์ด์ ํธ๊ฐ ํ๋ฐฉํ๋ ์ฑ๊ฒฉ์ ๋ฐ๋ผ ๋ชฉ์๋ฆฌ์ ํผ์น๋ ์๋์ง๊ฐ ๊ฐ๊ธฐ ๋ค๋ฅด๋ค๊ณ ์ธ์งํ๋ค. ๋ํ, ์์ด์ ํธ์ ์ฑ๊ฒฉ์ ํน์ฑ์ ๊ธฐ๋ฐํด ์ธ์์ ์ผ๋ก ๊ธฐ๊ณ์ ๋จ์ฑ ํน์ ์ฌ์ฑ์ ์ ๋๋ฅผ ํฌ์ํ์ฌ ์ธ์ํ๋ ์์ธํ ํ์์ด ๊ด์ฐฐ๋์๋ค.
์ฅ์
์ฅ/๋จ๊ธฐ ์ข ๋จ์ ์ฐ๊ตฌ ์ฒด๊ณ: ๋จ์ ๋จํ์ฑ ์คํ์ค ์ธํฐ๋์ ์ ๊ทธ์น์ง ์๊ณ , 12์ผ๊ฐ์ ๊ฐ์ ๋ด ๋ฐฐํฌ(in-home deployment)๋ฅผ ํตํด ๊ณ ๋ น์ธต ์ฌ์ฉ์์ AI ์์ด์ ํธ ๊ฐ์ ๊ด๊ณ ํ์ฑ ๊ณผ์ ๊ณผ ์๊ฐ ๊ฒฝ๊ณผ์ ๋ฐ๋ฅธ ์ ํธ๋ ๋ฐ์ (flip) ํ์์ ์ ๋ฐํ๊ฒ ์ถ์ ํ๋ค.
๊ธฐ์ต ๋ณด์กด ๊ธฐ๋ฐ ๋ํ ๋ชจ๋ธ๋ง: ๋งค ๋ํ ์ธ์
์ ์ปจํ
์คํธ๋ฅผ ๊ตฌ์กฐํํ์ฌ AI์ ๋ฉ๋ชจ๋ฆฌ์ ๋์ ์ฃผ์
ํจ์ผ๋ก์จ, ๊ณ ๋ น์ธต ์ฌ์ฉ์๊ฐ "๋ด ๊ณผ๊ฑฐ ์ด์ผ๊ธฐ๋ฅผ ๊ธฐ์ตํ๊ณ ํ์ ์ง๋ฌธ์ ๋์ง๋ค"๋ ๊น์ ์ ์์ ์ ๋๊ฐ๊ณผ ์ฐ์์ฑ ์๋ ๋ํ ๋ง์กฑ๋๋ฅผ ๋ฌ์ฑํ๋ค.
ํ๊ณ์
์ ํ๋ ํ๋ณธ ํฌ๊ธฐ(N=5, N=3): ๊น์ด ์๋ ์ง์ ๋ถ์๊ณผ ์ผ๊ธฐ ๊ธฐ๋ก ์ฐ๊ตฌ(diary study)๋ฅผ ์ํด ์ฐธ๊ฐ์๋ฅผ ๋ฐ์ฐฉ ๊ด๋ฆฌํ์ผ๋, ์ ๋ฐ์ ์ธ ๊ณ ๋ น์ธต ์ธ๊ตฌ ํต๊ณํ์ ํน์ฑ์ผ๋ก ์คํ ๊ฒฐ๊ณผ๋ฅผ ๋๊ท๋ชจ๋ก ์ผ๋ฐํํ๊ธฐ์๋ ์ ๋์ ํ๋ณธ์ ํฌ๊ธฐ๊ฐ ๋ค์ ์ ํ์ ์ด๋ค.
ํ
์คํธ-์์ฑ ๋ณํ(TTS) ์ธํฐํ์ด์ค์ ๊ณ ์ ์ฑ: ํ๋กฌํํธ(Pยฒ) ์ ์ด๋ก ๋ํ์ ๋ด์ฉ์ ์ฑ๊ฒฉ์ ์๋ฒฝํ ๋ค๋ณํํ์ผ๋, ๋ชฉ์๋ฆฌ ํค์ด๋ ๋ฐํ ์๋ ๊ฐ์ ์ฒญ๊ฐ์ paralinguistic cues๊ฐ ํ๋์ ๊ณ ์ ๋ ํฉ์ฑ ์์ฑ์ผ๋ก ์ถ๋ ฅ๋์ด ์ฑ๊ฒฉ์ ์ธํ์ ์๋๊ฐ์ 100% ๊ทน๋ํํ์ง๋ ๋ชปํ๋ค.
ํฅํ๊ณผ์
์ฑ๊ฒฉ-์์ฑ ๋์ ์ฑํฌ๋ก๋์ด์ฆ๋ ์์ง: LLM์ด ์์ฑํ๋ ์ฑ๊ฒฉ ํ๋กฌํํธ์ ํค์ค๋งค๋์ ์ค์๊ฐ์ผ๋ก ์ฐ๋๋์ด, ์ธํฅํ ์ฑ๊ฒฉ์ผ ๋๋ ๋ฐํ ์๋๊ฐ ๋นจ๋ผ์ง๊ณ ์ต์์ด ๋ค์ด๋ด๋ฏนํด์ง๋ฉฐ, ์ฑ์คํ์ผ ๋๋ ์ฐจ๋ถํ๊ณ ์ผ์ ํ ํผ์น๋ก ํค์ด ์๋ ๋ณ์กฐ๋๋ ์ฑ๊ฒฉ-TTS ํตํฉ ํ์ดํ๋ผ์ธ์ ๊ตฌ์ถํ ์์ ์ด๋ค.
์ฌ์ฉ์ ์ํ ๋ง์ถคํ ๊ฐ๋ณ ์ฑ๊ฒฉ ์ ์ด ์๊ณ ๋ฆฌ์ฆ: ์์ด์ ํธ์ ์ฑ๊ฒฉ์ ํ๋๋ก ๊ณ ์ ํ๋ ๋์ , ๊ณ ๋ น์ธต ์ฌ์ฉ์์ ํ์ฌ ๋น์ผ ๊ฐ์ ์ํ(์ฌํ, ์ธ๋ก์, ํ๊ธฐ์ฐธ ๋ฑ)๋ฅผ ๋ํ ์๋์์ ์ค์๊ฐ ๊ฐ์งํ์ฌ, ์๋ก๊ฐ ํ์ํ ๋๋ ์นํํ์ผ๋ก, ๋ฌด๋ฃํจ์ ๋ฌ๋๊ณ ์ถ์ ๋๋ ์ธํฅํ์ผ๋ก ์ฑ๊ฒฉ์ ๋๋๋ฅผ ๋์ ์ผ๋ก ๋ธ๋ ๋ฉํ๋ ๊ฐ๋ณํ CA ์๊ณ ๋ฆฌ์ฆ์ ๊ณ ๋ํํ ๊ณํ์ด๋ค.
Md Aashikur Rahman Azim, Zihao Su, Seongkook Heo
Your Hands Can Tell: Detecting Redirected Hand Movements in Virtual Reality
(Abstract) In-air hand interactions are prevalent in Virtual Reality (VR), and prior studies have shown that manipulating the visual movement of the hand to be different from the actual hand movement, i.e., hand redirection, could create a more immersive and engaging VR experience. However, this manipulation risks degrading task performance and, if maliciously applied, poses a threat to user safety. Such manipulations may arise from VR applications developed with intentional or inadvertent perceptual manipulations that yield harmful outcomes. We advocate for a userโs prerogative to be informed of any such potential manipulations before application usage. To address this, our study introduces an Autoencoder-based anomaly detection technique that leverages usersโ inherent hand movements to identify hand redirection, thereby preserving the integrity of application use. Our model is trained on regular (i.e., non-manipulated) hand movement patterns and employs a stochastic thresholding approach for anomaly detection. We validated our method through a technical evaluation involving 21 participants engaged in reaching tasks under manipulated and non-manipulated scenarios. The results demonstrated a high accuracy of hand redirection detection at 93.7%, with an F1-score of 93.9%.
(Introduction)ย Over the past few years, Virtual Reality (VR) has gained a significant rise in popularity, driven by advancements in high-fidelity 3D visual displays and 6-degree-of-freedom (DOF) tracking technologies. These developments have expanded VRโs application across diverse fields, including entertainment, education, and work-related tools [4, 28, 47, 54, 56]. VRโs unique capability to immerse users in a completely virtual environment, distinct from the physical world, offers unprecedented opportunities for enhancing user experiences. One of the ways VR utilizes to enhance user experiences is perception manipulation, which leverages visual dominance to manipulate usersโ perception by subtly altering the perceived reality [19]. Hand redirection is one of perception manipulation techniques to manipulate perception, such as diverging the visual representation of hand movements from their actual trajectory [5, 13, 31], creating illusions of weight [45], and altering perceived object size [2], have been extensively studied. Additionally, the use of dynamic movement gains to improve ergonomic interactions within VR environments has also been explored [55].
Despite their promise, these dynamic manipulation techniques may inadvertently impact the precision of skilled motor movements, which are developed through extensive repetition. Humans rely on both visual and kinesthetic feedback to refine and optimize their movements, learning to adjust based on the disparity between intended and actual movements [33, 38, 46, 50]. The acquisition of such motor skills and the development of muscle memory is crucial for performing tasks with increased speed, consistency, and stability, with minimal cognitive load [34]. However, the dynamic nature of perception manipulation, which alters the mappings between virtual and real hand movements, poses challenges to the development of such automaticity. Moreover, there are concerns that perception manipulation could be maliciously exploited to harm users, with potential adversaries including VR developers who might intentionally or unintentionally introduce harmful manipulations [51]. Instances of such exploitation have been documented, including manipulation of VR safety boundaries to control user movement without their awareness [10]. These studies show that perception manipulation can be harmful and pose significant risks when used maliciously. Such manipulations may degrade performance, compromise user safety, or exploit vulnerabilities. Detecting these manipulations is challenging, especially when developers do not disclose their code or inform users about the applicationโs manipulation state. To protect users from such manipulations, it is crucial to detect such manipulations through alternative means. In this work, we propose a semi-supervised anomaly detection approach that detects subtle changes in the userโs behavior during reaching tasks.
This work introduces a novel approach to detecting hand redirections, which is a type of perception manipulation in VR. We leverage natural hand movements within the theoretical framework of Woodworthโs two-component model of goal-directed movements, established in 1899 [57]. According to this model, goal-directed actions comprise an initial ballistic phase, directing the limb toward the target, followed by a corrective phase that adjusts for trajectory errors using visual feedback. Our observations indicate that the speed of usersโ corrective movements varies in response to perception manipulation, a variation not present under normal conditions. This phenomenon occurs when perception manipulation is applied in a particular direction, causing users to adjust their hand movements in the opposite direction. One can easily notice this adjustment, which is evident in the hand movement trajectories shown in Figure 1 and Figure 5. This observation forms the basis of our hypothesis that the characteristics of hand movements during corrective phases, such as speed, acceleration, and jerk, can serve as indicators to differentiate between manipulated and normal movement patterns.
To test our hypothesis, we developed a semi-supervised machine learning model based on an Autoencoder (AE) [6, 15] for anomaly detection. The key aspect of using Autoencoder is that we do not need to train the model on manipulated data. Instead, the model is trained solely on the userโs normal movement data, identifying anomalies in movement patterns during testing. In other words, by leveraging on the normal movement data, the model distinguishes between manipulated and non-manipulated instances by learning to reconstruct normal movement patterns, with the premise that anomaliesโmanipulated movementsโwill exhibit higher reconstruction errors [6, 62]. This approach leverages the assumption that normal instances are more accurately reconstructed from learned representations, while anomalies deviate significantly, resulting in higher reconstruction errors. The determination of an anomaly is contingent upon a reconstruction error exceeding a predetermined threshold. To optimize this threshold, we conducted multiple evaluations of our model, selecting the optimal threshold based on the differentiation between normal and manipulated data.
The efficacy of our approach was validated through a user study involving 21 participants, where we collected hand movement data in VR under both normal (non-manipulated) and manipulated conditions. Our study design incorporated orientation manipulation at two levels (ยฑ10 ยฐ ) to simulate manipulated movement conditions. Data collection was structured into three sessions: an initial session to establish baseline hand movements (S1), a second session to collect data under manipulated conditions (S2), and a final session to gather additional baseline data (S3). The model was trained in a semi-supervised manner using only S1 data to learn regular hand movement patterns, with S2 and S3 data serving to evaluate the modelโs ability to accurately detect manipulated instances as anomalies and baseline instances as normal. Our findings, derived from a within-subject analysis employing a 70-30 train-test split, demonstrate that our model achieves an anomaly detection accuracy of 97.7% (F1-score: 98.1%). Further validation through leave-one-participant-out (LOPO) cross-validation confirmed the robustness of our model in detecting anomalies across new users, with an accuracy of 93.7% (F1-score: 93.9%).
The contributions of this paper are twofold:
Introduction of a novel deep learning approach for detecting the presence of hand redirection in VR from userโs hand movement behavior. To our knowledge, this is the first method to automatically detect the hand redirection in VR from the userโs behavior data.
The dataset of manipulated and non-manipulated hand movement data1 from VR tasks, addressing the absence of publicly available datasets for hand redirection research.
(Conclusion) This study introduced a semi-supervised AE model for detecting hand movement manipulations in VR environments, leveraging velocity and acceleration profiles from usersโ hand movements. Through a comprehensive analysis, employing both within-subjects and LOPO approaches, the model demonstrated high accuracy (93.7%) in identifying anomalies, underpinned by the careful selection of an optimal anomaly threshold. Despite facing limitations such as the absence of SSs in certain instances and the studyโs duration, the research highlights the modelโs potential to enhance VR security by preemptively identifying and warning users of manipulative applications. Future investigations will aim to broaden the modelโs applicability by incorporating varied manipulation types and extending the observational period. This work lays a foundational step towards developing robust, user-centric anomaly detection systems in VR, promising a safer interactive experience for users across diverse VR applications.
๊ฐ์ํ์ค(VR) ํ๊ฒฝ์์ ์ ์๊ณก ๊ธฐ์ ์ ์ฅ์ ๊ณผ ํ๊ณ
์ ์๊ณก(hand redirection) ๊ธฐ์ ์ ์ฅ์ : VR ๋ด์์ ์์ ํ๊ณต์ ์์ง์ด๋ bare-hand ์ธํฐ๋์ ์, ์๊ฐ์ ์ผ๋ก ๋ณด์ด๋ ๊ฐ์ ์์ ์์น๋ ๊ถค์ ์ ์ค์ ์์ ๋ฌผ๋ฆฌ์ ์์ง์๊ณผ ๋ค๋ฅด๊ฒ ๋ณ์กฐํจ์ผ๋ก์จ ํ์ ๋ ๊ณต๊ฐ ๋ด์์ ๋ ๋ชฐ์ ๊ฐ ์๊ณ ๋ชฐ์ ํ(immersive)์ธ ์ฌ์ฉ์ ๊ฒฝํ์ ์์ฑํด ๋ผ ์ ์๋ค.
์ฌ์ฉ์ ์์ ๋ฐ ์์ ์ ๋ฐ๋ ์ ํด ์ํ: ํ์ง๋ง ๋ฏธ์ธํ ์๊ฐ-์ด๋ ๋ถ์ผ์น(visual-motor mismatch) ์กฐ์์ ์ค๋ ๋ฐ๋ณต์ ํตํด ๋์ ๊ทผ์ก์ ํ์ฑ๋ ์ฌ์ฉ์์ ์ ๋ฐํ ์ด๋ ๋ฅ๋ ฅ(motor skills)๊ณผ ๊ณ ์ ๊ฐ๊ฐ์ ๊ต๋ํ์ฌ ์์ ์ ํ๋๋ฅผ ๋จ์ด๋จ๋ฆด ์ ์๋ค.
์ ์์ ๋ณ์กฐ ๋ฐ ์์ ๊ฒฝ๊ณ ์กฐ์ ์ํ: ๋ง์ฝ, VR ์ดํ๋ฆฌ์ผ์ด์ ๊ฐ๋ฐ์๋ ์ธ๋ถ ์ ์ฑ ๊ณต๊ฒฉ์๊ฐ ์ฌ์ฉ์๊ฐ ์ธ์งํ์ง ๋ชปํ๋ ์ญ์น ๋ฏธ๋ง(sub-threshold) ์์ค์ ๋ฏธ์ธํ ์ธ์ง ์กฐ์์ ์ ์์ ์ผ๋ก ์ฃผ์ ํ ๊ฒฝ์ฐ, ์ฌ์ฉ์์ ์ ์ฒด ์์ง์์ ์์๋ก ๊ฐ์ ํต์ (์, human joystick attack)ํ๊ฑฐ๋, ๊ฐ์ ๊ฒฝ๊ณ์ ์ ์๊ณกํ์ฌ ๊ฐ๊ตฌ ์ถฉ๋ ๋ฑ ์ฌ๊ฐํ ๋ฌผ๋ฆฌ์ ์์ ์ฌ๊ณ ๋ฅผ ์ ๋ฐํ ์ ์๋ค.
๋ถํฌ๋ช
ํ ๋ธ๋๋ฐ์ค ์ดํ๋ฆฌ์ผ์ด์
์ ๋ณด์์ ํ๊ณ: third-party ์ดํ๋ฆฌ์ผ์ด์
์ด ์์ค์ฝ๋๋ฅผ ๊ณต๊ฐํ์ง ์๊ฑฐ๋, ํ๋ซํผ ๋ด๋ถ์ ์๊ณก ์๊ณ ๋ฆฌ์ฆ์ ์๋ฐํ ํฌํจํ๊ณ ์์ ๊ฒฝ์ฐ, ์์คํ
๋ ๋ฒจ์์ ์ด๋ฌํ ๋น์ ์์ ์กฐ์ ์ฌ๋ถ๋ฅผ ์ฌ์ ์ ๊ฐ์งํ๊ณ ์ฌ์ฉ์์ ์ ์ฒด์ ์์ ๊ณผ ํ๋ผ์ด๋ฒ์๋ฅผ ๋ฅ๋์ ์ผ๋ก ๋ณดํธํ ๋ณด์ ์ฅ์น๊ฐ ์ ๋ฌดํ๋ค.
Your Hands Can Tell
VR ํ๊ฒฝ ๋ด์์ ์ดํ๋ฆฌ์ผ์ด์ ์ด ์์ค ์ฝ๋๋ฅผ ์จ๊ธฐ๋๋ผ๋ ์ฌ์ฉ์์ ๋ฏธ์ธํ ์ ์์ง์ ์ญํ(hand movement dynamics) ๋ณํ๋ง์ ๋ถ์ํ์ฌ, ๋ฐฐํ์์ ์๋ฐํ๊ฒ ๊ฐํด์ง๋ ์ ์๊ณก(redirection) ์กฐ์์ ์ค์๊ฐ์ผ๋ก ๊ฐ์งํด ๋ด๋ ์คํ ์ธ์ฝ๋(autoencoder) ๊ธฐ๋ฐ์ ํ์ ์ ์ธ VR ๋ณด์ ๋ฐ ์ด์ ํ์ง(anomaly detection) ์์คํ ์ด๋ค.
์ฐ๋์์ค(Woodworth)์ 2๋จ๊ณ ์ด๋ ๋ชจ๋ธ ๊ธฐ๋ฐ ํผ์ฒ ์ถ์ถ: ์ธ๊ฐ์ ๋ชฉํ ์งํฅ์ ๋๋ฌ ๋์์ด ์ด๊ธฐ ๊ณ ์ ํ๋ ๋จ๊ณ(PS, Primary Submovement)์ ์๊ฐ์ ํผ๋๋ฐฑ์ ํตํด ๋ชฉํ์ ์ ์์ฐฉํ๋ ๋ฏธ์ธ ๋ณด์ ๋จ๊ณ(SS, Secondary Submovement)๋ก ๋๋๋ค๋ ์ ์ ์ฐฉ์ํ๋ค. ์๊ฐ์ ์๊ณก์ด ์ฃผ์ด์ง๋ฉด ์ฌ์ฉ์๊ฐ ๋ฌด์์์ ์ผ๋ก ๋ฐ๋ ๋ฐฉํฅ ๋ณด์ ์ ์ด๋ฅผ ์ํํ๋๋ผ SS ๋จ๊ณ์ ์๋(velocity)์ ๊ฐ์๋(acceleration) ํ๋กํ์ ๋ฏธ์ธํ ์ด์ ํํ์ด ์๊ธด๋ค๋ ์ ์ ๊ท๋ช ํ์ฌ ํต์ฌ ํน์ง(feature)์ผ๋ก ์ถ์ถํ๋ค.
LSTM ๊ธฐ๋ฐ ์คํ ์ธ์ฝ๋(AE) ๋ฐ์ง๋ ํ์ต(semi-supervised): ์กฐ์๋ ๊ณต๊ฒฉ ์ธ์คํด์ค ๋ฐ์ดํฐ๋ฅผ ์ผ์ผ์ด ํ์ตํ ํ์ ์์ด, ์กฐ์์ด ์๋ ์ํ์์ ์์ง๋ ์ฌ์ฉ์์ ์ ์์ ์ธ(Normal) ์ ์์ง์ ํจํด๋ง์ ์ ๊ฒฝ๋ง์ ํ์ต์ํจ๋ค. ์ ์ ๋์ ์ํ์ค๋ ์์ถ ํ ์๋ฒฝํ ๋ณต์(๋ฎ์ ๋ณต์ ์ค์ฐจ)ํ์ง๋ง ์๊ณก ์กฐ์์ด ๊ฐํด์ง ๋์์ ๋ณต์ ๋ฅ๋ ฅ์ด ํ์ ํ ๋จ์ด์ ธ ์ ์ฌ ๊ณต๊ฐ(latent space) ์ค์ฐจ๊ฐ ๊ธ๊ฒฉํ ์ปค์ง๋ ํน์ฑ์ ์ด์ฉํด ์ด์์น(anomaly)๋ก ๋ถ๋ฅํ๋ค.
์คํ ์บ์คํฑ ์๊ณ๊ฐ ๋ฐ ๋ฐ์ดํฐ์
๊ฒ์ฆ(N=21): ์์คํ
๊ฒ์ฆ์ ์ํด 21๋ช
์ ์ฐธ๊ฐ์๋ฅผ ๋์์ผ๋ก VR ํ๊ฒฝ์์ ์ด 3๊ฐ ์ธ์
(๊ธฐ๋ณธ ์ธก์ S1, ยฑ 10ยบ ๋ฐฉํฅ ์ ์๊ณก ์ฃผ์
S2, ๋ณต๊ท ์ธก์ S3)์ ๋๋ฌ(reaching) ํ์คํฌ๋ฅผ ์ํํ๊ฒ ํ์ฌ ์ด 4,704๊ฐ์ ์ ์ baseline ๋ฐ์ดํฐ์ 1,680๊ฐ์ ์กฐ์ ๋ฐ์ดํฐ๋ฅผ ์์ง ๋ฐ ๊ฒ์ฆํ๋ค.
์ฐ๊ตฌ ๊ฒฐ๊ณผ ๋ฐ ๋ณด์ํ์ ์์
๊ทน๋๋ก ๋์ ์์ค์ ์กฐ์ ํ์ง ์ ํ๋ ๋ฌ์ฑ: ํผํ์ ๋ด(within-subjects) ๋ถ์ ํ๊ฒฝ(70-30 train-test split)์์ 97.7%์ ์ต๊ณ ์์ค ์ด์ ํ์ง ์ ํ๋(F1-score: 98.1%)๋ฅผ ๊ธฐ๋กํ์ต๋๋ค. ํนํ ์์ ํ ์๋ก์ด ์ฌ์ฉ์๋ฅผ ๋์์ผ๋ก ๋ชจ๋ธ์ ํ๊ฐํ๋ ํผํ์ ๊ฐ(LOPO,Leave-One-Participant-Out) ๊ต์ฐจ ๊ฒ์ฆ์์๋ 93.7%์ ๋์ ์ผ๋ฐํ ์ ํ๋(F1-score: 93.9%)๋ฅผ ํ๋ณดํ์ฌ ๋ชจ๋ธ์ ๊ฐ๋ ฅํ ๊ฐ๊ฑดํจ(Robustness)์ ์ฆ๋ช ํ๋ค.
์ผ๊ด๋ ๋ฒ์ฉ ์๊ณ๊ฐ(๐=6)์ ์ ํจ์ฑ ๊ฒ์ฆ: ์ํ์ ์ต์ ํ ๋ฉ์ปค๋์ฆ์ ํตํด ๋์ถํ ์ด์ ํ์ง ์๊ณ์น ์์น(๐=6)๊ฐ ๊ฐ์ธ๋ณ ๋ง์ถคํ ์ธํ ๋ฟ๋ง ์๋๋ผ, ์์คํ ์ด ์ฒ์ ๋ณด๋ ์๋ก์ด ์ฌ์ฉ์๋ฅผ ๋์ ํ LOPO ๋ถ์ ํ๊ฒฝ์์๋ ์์ ํ ์ผ๊ด๋๊ฒ ์ ์ฉ ๊ฐ๋ฅํจ์ ๋ฐ๊ฒฌํ๋ค. ์ด๋ ๋ณ๋์ ์ฌํ์ต ์์ด ์ค์๊ฐ ์จ-์คํ๋ผ์ธ ์์คํ (online deployment)์ ์ฆ์ ๋์ ํ ์ ์์์ ์ ๋์ ์ผ๋ก ๋ท๋ฐ์นจํ๋ค.
์ ์ ์ ๊ฐ์ํ์ค ๋ณด์ ์ธํ๋ผ(framework) ์ ์: ๋ณธ ๊ธฐ์ ์ VR ์ด์์ฒด์ ์ํคํ
์ฒ๋ ์ฑ์คํ ์ด ๋ณด์ ๊ฒ์ ์ธํ๋ผ์ ํตํฉํ ๊ฒฝ์ฐ, ์ฌ์ฉ์๊ฐ ์๋ก์ด ์ฑ์ ์คํํ์ฌ ๋ช ๋ฒ์ ์์ฐ์ค๋ฌ์ด ์ ์์ง์ ์ํธ์์ฉ์ ์ํํ๋ ๊ณผ์ ์์ ์์คํ
์ด ๋ฐฐํ์ ์
์์ ์ธ ์ธ์ง ์กฐ์์ ์๋์ผ๋ก ๊ฐ์งํ๊ณ ์ํ ๊ฒฝ๊ณ ๋ฅผ ๋ณด๋ผ ์ ์๋ ์ฌ์ฉ์ ์ค์ฌ์ ๋ฅ๋ํ VR ๋ณด์ ํ ๋๋ฅผ ๋ง๋ จํ๋ค.
์ฅ์
์ ๋ก ๊ณต๊ฒฉ ๋ฐ์ดํฐ ํ๋ จ: ์ค์ง ์ ์ ๋ฐ์ดํฐ(normal dataset)๋ง์ ๊ธฐ๋ฐ์ผ๋ก ์คํ ์ธ์ฝ๋๋ฅผ ํ์ต์ํค๋ ๋ฐ์ง๋ ํ์ต ๋ฐฉ์์ ์ฑํํ์ฌ, ํฅํ ์ด๋ค ๋ณํ๋ ํํ๋ ๋ฏธ์ง์ ์ ์๊ณก ๊ณต๊ฒฉ(zero-day redirection attacks)์ด ๋ค์ด์ค๋๋ผ๋ ์ฌ์ ์ ์ ์ฐํ๊ฒ ์ก์๋ผ ์ ์๋ค.
๋
๋ฆฝ์ ๋ธ๋๋ฐ์ค ์คํผ๋ ์ด์
: ๋์ ์ดํ๋ฆฌ์ผ์ด์
๋ด๋ถ์ ๋ ๋๋ง ์์ค ์ฝ๋๋ ์
ฐ์ด๋ ํจ์ ๊ตฌ์กฐ๋ฅผ ๋ฏ์ด๋ณผ ํ์ ์์ด, ํค๋๋ง์ดํธ ๋์คํ๋ ์ด(HMD)๊ฐ ํธ๋ํนํ๋ ์ฌ์ฉ์์ ์์ ์์ ์ ์ขํ ์ ๋ณด(kinematic data)์๋ง ์์กดํ๋ฏ๋ก ํ๋ซํผ ๋ณด์ ๋ ์ด์ด๋ฅผ ํด์น์ง ์๊ณ ๋ฐฑ๊ทธ๋ผ์ด๋์์ ๊ตฌ๋๋๋ค.
ํ๊ณ์
์กฐ์ ํฌ๊ธฐ(magnitude)์ ๋ฐ๋ฅธ ํ์ง ํด์๋ ๋ณ๋: ์ ์๊ณก์ด ๊ฐํด์ง๋ ๊ฐ๋๊ฐ ยฑ 10ยบ ์์ค์ผ ๋๋ ์ด๋ ์ญํ ๋ณํ๊ฐ ๋๋ ทํ์ฌ ์๋ฒฝํ ๊ฐ์งํด ๋ด์ง๋ง, ๋ง์ฝ ๊ณต๊ฒฉ์๊ฐ ์ฌ์ฉ์์ ๊ณ ์ ๊ฐ๊ฐ ์ญ์น(detection threshold)๋ณด๋ค๋ ํจ์ฌ ๋ฏธ๋ฏธํ๊ณ ๊ทน์ํ ์์ค(์, 1ยบ ~ 3ยบ ๋ฏธ๋ง)์ผ๋ก ์ค๋ ์๊ฐ์ ๊ฑธ์ณ ์์ฃผ ์ฒ์ฒํ ์๊ณก์ ๊ฐํ ๊ฒฝ์ฐ ์คํ ์ธ์ฝ๋์ ๋ณต์ ์ค์ฐจ๊ฐ ์๊ณ์น(๐=6)๋ฅผ ์ฆ๊ฐ ๋์ง ๋ชปํด ์ด๊ธฐ ๊ฐ์ง ๋ ์ดํด์๊ฐ ๊ธธ์ด์ง ์ ์๋ค.
๋๋ฌ ๋์(reaching task) ์ค์ฌ์ ํผ์ฒ ์ ํ: ๋ณธ ๋ชจ๋ธ์ ์ฐ๋์์ค์ ๋ชฉํ ์งํฅ์ ํ๋ ์ด๋ ๋ฒ์น์ ๊ฐํ๊ฒ ๋ฐ์ธ๋ฉ๋์ด ํผ์ฒ๋ฅผ ์ถ์ถํ๋ฏ๋ก, ์ ํํ๋ ๋๋ฌ ๋์์ด ์๋ VR ๊ฒ์ ๋ด์์์ ๋ณต์กํ ์์ ํ ๋๋ก์, 3D ๊ฐ์ ๊ฐ์ฒด ์กฐ๋ฆฝ, ์์ ์ ์ค์ฒ ๋ฑ ์ฐ์์ ์ด๊ณ ๋ถ๊ท์นํ ์์ ๋์(free-form manipulation) ํ๊ฒฝ์์๋ ์ด์ ํจํด์ ์ ๋ฐํ๊ฒ ๋ถ๋ฆฌํด๋ด๊ธฐ ์ํ ํผ์ฒ ํ์ดํ๋ผ์ธ์ ์ถ๊ฐ ๊ณ ๋ํ๊ฐ ํ์ํ๋ค.
ํฅํ๊ณผ์
์์ ํํ ์ธํฐ๋์ ๋ฐ ์ํ์ค ์ต์ ํ: ํน์ ํ์คํฌ ์๋๋ฆฌ์ค์ ์ข ์๋์ง ์๊ณ ์ฌ์ฉ์๊ฐ ๊ฐ์ ํ์ค ๋ด์์ ํํ๋ ๋ชจ๋ ๋ถ๊ท์นํ ์๋๋ฆผ(continuous free-form movements) ๋ฐ์ดํฐ ํ๋ฆ ์์์๋ ์ ์๊ณผ ์๊ณก ์ํ๋ฅผ ์ค์๊ฐ ๊ฐ๋ ์ค์ฐจ ์๋์ฐ๋ก ๊ตฌ๋ณํด๋ผ ์ ์๋๋ก ์๊ณต๊ฐ ์ดํ ์ ๋ฉ์ปค๋์ฆ(spatial-temporal attention)์ ์คํ ์ธ์ฝ๋์ ๊ฒฐํฉํ ์์ ์ด๋ค.
๋ค์ค ์ธ์ง ๋ณ์กฐ ํตํฉ ํ์ง ์ธํ๋ผ ํ์ฅ: ์์ ๋ฌผ๋ฆฌ์ ์์น ์๊ณก์ ๋์ด, ์์ผ๊ฐ์ ๊ฐ์ ๋ก ๋๋ฆฌ๋ ์์ ์๊ณก(gain manipulation), ๊ฐ์ ๊ณต๊ฐ์ ํฌ๊ธฐ๋ฅผ ๋ฏธ์ธํ๊ฒ ์์ถํ๋ ๊ณต๊ฐ ๊ฐ์ํ ์๊ณก(redirected walking), ๊ทธ๋ฆฌ๊ณ ํ๋ ์ ๋ ์ดํธ๋ฅผ ์ธ์์ ์ผ๋ก ๋ฏธ์ธ ๋ณ์กฐํ๋ ์ ์ฌ์ ์ง์ฐ ๊ณต๊ฒฉ ๋ฑ VR ํ๊ฒฝ ์ ์ฒด๋ฅผ ์ํํ๋ ๋ค์ํ ๋ค์ฐจ์ ์ธ์ง ์กฐ์ ๊ณต๊ฒฉ์ ํตํฉ ๋ชจ๋ํฐ๋งํ๊ณ ๋๋ฒ๊น ํ ์ ์๋ ์ข ํฉ์ ์ธ ์จ๋๋ฐ์ด์ค VR ๋ฉด์ญ ์์คํ ์ ๊ตฌ์ถํ ๊ณํ์ด๋ค.
Buyoung Mun, Junsu Lee, Jiseong Kim, Seongkook Heo, Jaeyeon Lee
Diversifying Grain-Based Compliance Illusion by Varying Base Compliance
(Abstract) Grain-based compliance illusion mimics the mechanical vibrations that occur when a compliant object deforms with grain-like, short (โผ15 ms) impulse-response vibrations. Previous work has demonstrated its robust effect on various types of devices. However, the impact of the deviceโs inherent compliance (i.e., base compliance) on perceived compliance remains unclear. This paper investigates the influence of base compliance on the perception of illusory compliance through three psychophysical experiments. The results show that (1) the compliance illusion remained effective with base compliance, (2) the description of compliance was affected by both illusory and base compliance, and (3) it is possible to render the compliance with the same magnitude but multiple different feelings.
(Introduction) It is crucial for people to feel the mechanical properties of an object, such as its softness and elasticity, as well as its behavior, to conduct dexterous manipulation in the real world, as it allows people to understand the material and the status of the object. For humancomputer interaction, haptic interfaces render these properties to allow users to manipulate 3D models effectively [39], precisely conduct teleportation [19], and increase the immersiveness of the virtual experience [32]. However, rendering such mechanical properties often comes at the cost of an expensive, large, and grounded device that uses mechanical or magnetic mechanisms to alter the physical properties of the haptic interface. These constraints limit their portability and usability in mobile scenarios.
To address this limitation, many illusion-based techniques have been developed to alter usersโ perception rather than the physical properties of the interface. One widely used method is the visuohaptic illusion, which leverages the visual dominance of human perception. By manipulating the visual representation of an object, these techniques can make users perceive changes in stiffness [2, 48], shape [1], and weight [38]. Although effective, these methods depend on usersโ visual attention and only work when the object is visually rendered and users are looking at the object. In contrast, tactile illusions create sensations of movement [6, 35] or deformation [11, 16, 21] without requiring structural changes in the object. These illusions rely on tactile actuators, such as vibrotactile mechanisms, to simulate different material properties through direct touch. This approach enables a more natural sensation, as users feel the object directly with their hands, akin to real-world interaction. Moreover, tactile illusions do not require users to look at the object, freeing their visual attention to focus on other objects or tasks.
Among tactile illusion techniques, grain-based compliance illusions have gained significant interest within the HCI community in recent years. These techniques are inspired by the psychophysical study conducted by Srinivasan and Lamotte [40], which demonstrated that participants could successfully discriminate between objects of varying stiffness using only tactile sensations. Building on this insight, researchers developed techniques utilizing grainlike short vibrations that mimic the tactile feedback caused by movement or deformation in response to force input. These studies revealed that such vibrations could create the illusion of compliance without actual physical movement or deformation [15, 16, 20, 21, 23]. Moreover, studies have shown that the properties of these illusions can be programmatically manipulated. For example, altering the virtual displacement threshold for triggering grain vibrations changes the perceived stiffness, while capping feedback for forces exceeding a predefined threshold restricts the sensation of further movement or deformation. This programmability has enabled grain-based compliance illusions to render tactile buttons with customizable force-displacement characteristics [22, 23] and virtual joints with varying movement constraints [16]. The effectiveness, adaptability, and simplicity of grain-based compliance techniquesโrequiring only a force sensor and a vibrotactile actuatorโmay explain their growing popularity in HCI research. These techniques have been demonstrated to enhance the precision [25] and realism [16, 41] of VR experiences, support high-fidelity tangible user interfaces [37], and facilitate rapid prototyping of tangible interfaces [44].
While grain-based illusion techniques have shown promising results in creating the illusion of compliance on a rigid surface enhancing the usability and user experience of interfaces, the setups used in these studies often involve an unclear parameter that may significantly influence the perception of complianceโthe rigidity of the so-called rigid object. Materials that we consider rigid, such as plastic or steel, have some level of compliance. Given that grain-based illusion methods utilize force as input, devices used in previous studies may have a different level of deformation during the study. For example, some of the previous studies used acrylic plates under the finger [15, 25], 3D printed blocks [23] and handles [16], and an unknown rigid material [21] that vary in different sizes. In these studies, while the surface itself may not locally deform around the finger pad to alter the pressure distribution in the area of contact to make it feel as soft [27, 34, 42], the deformation of the base structure may have caused kinesthetic compliance (i.e., device being bent). This reveals two key issues: first, the diverse effects of base material compliance on perceived compliance properties have been largely overlooked; and second, this makes designers and engineers seeking to use grain-based illusion techniques to properly account for critical design properties and understand the resulting illusion.
In this paper, we investigate the effect of base compliance on the perception of grain-based compliance illusion, including the range of perceived compliance, Just Noticeable Difference (JND), and the perceived properties of compliance through a series of psychophysical experiments. We implemented a prototype device with a load cell and a vibrotactile actuator that plays short bursts of vibrations (i.e., grain vibrations) per a certain level of force changes to create a grain-based compliance illusion. From the top, the device consisted of a rigid finger-resting plate with a vibrotactile actuator, an interchangeable base compliance block, and a force sensor. By changing the base compliance block, we could investigate three conditions with different levels of base compliance while keeping the illusory compliance and the surface condition consistent.
Our results showed how base compliance affects perceived compliance. First, the grain-based compliance illusion remained effective with the base compliance. The softer the base compliance is, the softer the device feels with the compliance illusion. We found that the overall range of compliance illusion shifts a certain amount depending on the base compliance (Figure 1). Second, the existence and the level of base compliance affect the description of perceived compliance. While the illusory compliance dominated, the existence and the level of base compliance made perceived difference. The illusory compliance with more compliant base material was perceived lighter, softer, deeper, and more granular compared to the hard base material. Thirdly, we found a new design opportunities of compliance created by base ร illusory compliance. Based on our finding in Experiments 1 and 3, by changing both base and illusory compliance, one can adjust both magnitude and feeling of compliance. In other words, one can create multiple feelings of compliance with the same magnitude of compliance. Finally, as a first investigation towards the unknown perceptual space created by base and illusory compliance, we provide design guidelines, baseline data in an interactive visualization, and areas to discover for future researchers.
The contributions of this paper are as follows:
The findings of three psychophysical experiments investigating the effect of base compliance on the perception of grain-based compliance illusion
The method of diversifying the range and the quality of perceived compliance by combining the base compliance and the illusory compliance
The findings on unknown perceptual spaces created by combining base and illusory complianceย
(Conclusion) With a prototype device with a rigid surface, interchangeable compliant blocks, and different levels of grain-based compliance illusion thresholds, we explored the effect of base compliance on the perception of grain-based compliance illusion, including the range of perceived compliance, JNDs, and the quality of compliance. The results showed that the grain-based compliance illusion worked robustly regardless of base compliance. Finally, we found that the existence of base material affects the quality of compliance using multidimensional scaling and adjective rating. We hope our findings foster future HCI research on rendering, augmenting, and quantifying grain-based compliance illusions.
๊ธฐ์กด ์ ์ฐ์ฑ ์ฐฉ์(compliance illusion) ์ฐ๊ตฌ์ ํ๊ณ
๊ธฐ๊ธฐ ๊ณ ์ ์ ๊ฐ์ฑ(rigidity) ๋งค๊ฐ๋ณ์ ๋๋ฝ: ๊ธฐ์กด ์ฐ๊ตฌ๋ค์ ๋ฑ๋ฑํ ํ๋ฉด ์์์ ์ง๋์ ํตํด ์ ์ฐ์ฑ ์ฐฉ์๋ฅผ ์ฑ๊ณต์ ์ผ๋ก ๊ตฌํํด๋์ผ๋, ์ ์ ์ํฌ๋ฆด์ด๋ ๊ธ์ ๋ฑ ์คํ์ ์ฌ์ฉ๋ ๊ธฐ๊ธฐ ์์ฒด์ ๋ฌผ๋ฆฌ์ ๋จ๋จํจ ์์ค(base compliance/rigidity)์ด ์ธ์งํ์ฑ์ ๋ฏธ์น๋ ์ํฅ์ ๊ฐ๊ณผํด ์๋ค.
๋์ผ ์ง๋ ๋์์ธ์ ํํ ์คํํธ๋ผ ํ๊ณ: ์ํํธ์จ์ด์ ์ผ๋ก ์ง๋ ํธ๋ฆฌ๊ฑฐ ์๊ณ๊ฐ(grain threshold)์ด๋ ์ง๋ ์งํญ์ ์กฐ์ ํ๋ ๋ฐฉ์๋ง์ผ๋ก๋ ๊ฐ์ ์ ์ฐ์ฑ์ ํฌ๊ธฐ(Magnitude) ์ธ์ ๊ฐ๊ฐ์ ์ด์ง๊ฐ(texture/feelings)์ ์คํํธ๋ผ์ ๋ํ๋ ๋ฐ ํ๊ณ๊ฐ ์กด์ฌํ๋ค.
๊ธฐ๊ธฐ ์ ์ฐ์ฑ๊ณผ ์ฐฉ์ ์ ์ฐ์ฑ์ ์ํธ์์ฉ ๋ฉ์ปค๋์ฆ ๋ถ์ฌ: ์จ์ด๋ฌ๋ธ ํ
ํฑ ์ฅ์น๋ ์ค๋งํธํฐ ํ๋ฉด์ฒ๋ผ ๋ค์ํ ๋ฌผ๋ฆฌ์ ๊ฐ์ฑ์ ๊ฐ์ง ์ค์ ํ๋ฉด ์์์ ์
์ ๊ธฐ๋ฐ ์ฐฉ์ ๊ธฐ์ ์ด ์๊ณก ์์ด ์๋ํ๋์ง, ๋ ์ ์ฐ์ฑ์ด ๊ฒฐํฉํ ๋์ ์ ์ฒด ์ด๊ฐ ์ธ์ง ๊ณต๊ฐ์ด ์ด๋ป๊ฒ ํ์ฑ๋๋์ง ๊ท๋ช
๋ ๋ฐ ์๋ค.
ํต์ฌ ๊ธฐ๋ฅ ๋ฐ ์ฐ๊ตฌ ๋ฐฉ๋ฒ๋ก
๊ฐ๋ณํ ๊ธฐ๋ณธ ์ ์ฐ์ฑ ํ ํฑ ํ๋กํ ํ์ ๊ฐ๋ฐ: ํ ์ผ์(load cell)์ ์ง๋ ์ก์ถ์์ดํฐ๊ฐ ์ฅ์ฐฉ๋ ๋จ๋จํ ์๊ฐ๋ฝ ๋ฐ์นจ๋ ์ฌ์ด์ ๊ต์ฒด ๊ฐ๋ฅํ ๊ธฐ๋ณธ ์ ์ฐ์ฑ ๋ธ๋ก(base compliance block)์ ์ฝ์ ํ ์ ์๋ ๊ธฐ๊ธฐ๋ฅผ ์ค๊ณํ๋ค. ๋ฌผ๋ฆฌ์ ๊ฐ์ฑ์ด ๋ค๋ฅธ ์ธ ๊ฐ์ง ์กฐ๊ฑด(soft: ecoflex 00-10, medium: dragonskin 10A, hard: ์ํฌ๋ฆด ํ๋ ์ดํธ)์ ๊ตฌ์ถํ๋ค.
์๋ฒ ๋๋ ์ ์ด ๊ธฐ๋ฐ ์ ์ ์ฐฉ์ ๋ชจ๋ธ ๊ตฌํ: ์ ์ฐ์ฑ ์ฐฉ์์ ์ญ์์ธ ๊ฐ์ ๋ณ์ ๋จ์์ธ illusory compliance [grain/N] ๋ชจ๋ธ์ ์ ์ฉํ์ฌ ์ฌ์ฉ์๊ฐ ๊ฐํ๋ ํ์ ๋ฐ๋ผ ~15ms์ ์งง์ ์ถฉ๊ฒฉ ์๋ต ์ง๋(grain vibration)์ ์์ฑํ์ฌ, ์ต๋ 4.0 grain/N ๋ฒ์ ๋ด์์ 5๊ฐ ๋จ๊ณ์ ์ฐฉ์ ์์ค์ ํ๋ก๊ทธ๋๋ฐ ์ ์ดํ๋ค.
3๋จ๊ณ ์ ์ ๋ฌผ๋ฆฌํ(psychophysical) ์คํ ์ค๊ณ: ์ฌ์ฉ์์ ์๊ฐ๋ฝ ์์ธ, ๋๋ฅด๋ ์๋(6.25 N/s), ์ต๋ ํ(5 N)์ ์๊ฐ ๊ฐ์ด๋๋ฅผ ํตํด ์๊ฒฉํ ํต์ ํ ์ฑ, (1) ์ ์ฐ์ฑ์ ํฌ๊ธฐ ์ถ์ (magnitude estimation), (2) ์ต์ ์๋ณ ์ฐจ์ด(JND) ์ธก์ , (3) ๋ค์ฐจ์ ์ฒ๋๋ฒ(MDS) ๋ฐ ํ์ฉ์ฌ ํ๊ฐ๋ฅผ ํตํ ์ธ์ง ๊ณต๊ฐ ๋ถ์์ ์ํํ๋ค.
์ฐ๊ตฌ ๊ฒฐ๊ณผ
๋ฌผ๋ฆฌ์ ๊ฐ์ฑ์ ๋ฐ๋ฅธ ์ฐฉ์ ๋ฒ์์ ํํ์ด๋(shift): ์ ์ ๊ธฐ๋ฐ ์ ์ฐ์ฑ ์ฐฉ์๋ ๊ธฐ๊ธฐ ๊ณ ์ ์ ๋จ๋จํจ๊ณผ ๊ด๊ณ์์ด ๋ชจ๋ ์กฐ๊ฑด์์ ๊ฐ๊ฑดํ๊ฒ(robust) ์๋ํ๋ค. ๋ค๋ง, ๋ฌผ๋ฆฌ์ ๋ฒ ์ด์ค ์ฌ์ง์ด ๋ถ๋๋ฌ์ธ์๋ก(soft) ์ฐฉ์ ํจ๊ณผ์ ์ ์ฒด์ ์ธ ์ธ์ง ๋ฒ์๊ฐ ๋ ๋ถ๋๋ฌ์ด ๋ฐฉํฅ์ผ๋ก ํํ์ด๋(shift)ํ๋ ๊ฑฐ๋์ ์ฆ๋ช ํ๋ค.
๊ณผ๋ํ๊ฒ ๋ถ๋๋ฌ์ด ๋ฒ ์ด์ค์ ๋ณ๋ณ๋ ฅ ์ ํ ํ์ธ: medium(10A)์ด๋ hard ์กฐ๊ฑด์ ๋นํด, ๋งค์ฐ ๋ถ๋๋ฌ์ด soft(00-10) ๋ฌผ๋ฆฌ ๋ฒ ์ด์ค ํ๊ฒฝ์์๋ ์ฌ์ฉ์๊ฐ ๊ฐ์ ์ ์ฐ์ฑ์ ๋จ๊ณ์ ๋ณํ๋ฅผ ๊ฐ์งํ๊ธฐ ์ํด ํจ์ฌ ๋ ํฐ ์ง๋ ์๊ณ๊ฐ ๋ณํ(๋ ํฐ JND)๊ฐ ์๊ตฌ๋์๋ค, i.e., ๋ค๋จ๊ณ ์ ์ฐ์ฑ์ ์ธ๋ฐํ๊ฒ ํํํ ๋ ๋๋ฌด ๋ง๋ํ ๊ธฐ์ด ์ฌ์ง์ ํผํด์ผ ํจ์ ๊ท๋ช ํ๋ค.
๋์ผํ ํฌ๊ธฐ, ๋ค๋ฅธ ๋๋(Base x Illusory Compliance)์ ์ค๊ณ ๊ธฐํ ๋ฐ๊ตด: ๋ฌผ๋ฆฌ์ ๊ฐ์ฑ๊ณผ ๊ฐ์ ์ง๋ ์ฐฉ์๋ฅผ ๊ฒฐํฉํ๋ฉด ๋์ผํ ์ ์ฐ์ฑ์ ํฌ๊ธฐ(magnitude)๋ฅผ ์ ์งํ๋ฉด์๋ ์์ ํ ๋ค๋ฅธ ์ง๊ฐ์ ๋ ๋๋งํ ์ ์๋ค. ๋ฒ ์ด์ค๊ฐ ๋ถ๋๋ฌ์ธ์๋ก ๊ฐ์ ์ ์ฐ์ฑ์ ๋ ๊ฐ๋ณ๊ณ (light), ๊น๊ณ (deep), ์ ์๊ฐ์ด ์์ผ๋ฉฐ(granular), ๊ณ ๋ฌด ๊ฐ์(rubbery) ๋๋์ ์ฃผ์ด ๊ฐ์ ์ง๊ฐ์ ๋ค์ฑ๋ก์ด ๋์์ธ ํํ๋ฒ์ ์ ์ํ๋ค.
ํ
ํฑ ์ดํ๋ฆฌ์ผ์ด์
์๋๋ฆฌ์ค ์ ์: ์นดํธ๋ฆฌ์ง ๊ตํํ ํธ๋ํฌ๋ ์ฅ์น, ๊ณ ์ ์ ์ฐ์ฑ ํ ์ ๋๋ ค ๋ฌผ๋ฆฌ ํ๋ฉด ๊ฐ์ฑ์ ์๋ ๋ณํํ๋ VR ์ปจํธ๋กค๋ฌ(haptic revolver ๊ตฌ์กฐ ์์ฉ), ๋จ์ผ ๋ฒ ์ด์ค ์์์ ๋์ ์ ์ฐ์ฑ ์์ญ์ ๋ง๋ค์ด๋ด๋ ์๊ฐ๋ฝ ์ฐฉ์ฉํ ์ฅ์น ๋ฑ์ ๊ตฌ์ฒด์ ์ธ ์ค๊ณ ๊ฐ์ด๋๋ผ์ธ์ ์ ๊ณตํ๋ค.
์ฅ์
์ด๊ฐ ํํ๋ ฅ์ ์ฐจ์ ํ์ฅ: ๋จ์ ์ ์ฐ์ฑ์ ๊ฐ๋(soft/hard) ๋ ๋๋ง์ ๋์ด, ๋ฌผ๋ฆฌ ๋ฒ ์ด์ค์ ๊ฐ์ ์๊ฐฑ์ด(grain)์ ๊ฒฐํฉ์ ํตํด ์ง๊ฐ์ ๊ฐ์ด(๊ฐ๋ฒผ์, ๊ณ ๋ฌด ๊ฐ์ ๋ฑ)์ ์์ญ๊น์ง ์ ์ดํ ์ ์๋ ๊ณ ์ฐจ์์ ํ ํฑ ๋์์ธ ๊ณต๊ฐ์ ํ์ฅํ๋ค.
์ ๋ฐํ ๊ฐ์ด๋๋ผ์ธ ์ ๊ณต: ์ ์ ๋ฌผ๋ฆฌํ ์คํ์ ํตํด ๋์ถ๋ JND ์์น์ MDS ๊ณต๊ฐ ๋งต์ ๋ฐํ์ผ๋ก, ํ๋์จ์ด ๋์์ด๋๋ค์ด ์จ์ด๋ฌ๋ธ์ด๋ ๋ชจ๋ฐ์ผ ๊ธฐ๊ธฐ ์ค๊ณ ์ ์ค์ธ์์ ํผํ ์ ์๋ ๋ฌผ๋ฆฌ์ ยท์ํํธ์จ์ด์ ์ต์ ์กฐํฉ ๊ฐ์ ์ ๋์ ์ผ๋ก ์ ์ํ๋ค.
ํ๊ณ์
๋์ ์ ๋ ฅ ์๋ ๋ณํ์ ๋ฐ๋ฅธ ์ธ์ง ๋ณ๋์ฑ: ๋ณธ ์คํ์ ์๊ฐ ๊ฐ์ด๋๋ฅผ ํตํด ์ฌ์ฉ์๊ฐ ๋๋ฅด๋ ์๋ ฅ ์๋(6.25 N/s)๋ฅผ ์ผ์ ํ๊ฒ ๊ฐ์ ํต์ ํ ํ๊ฒฝ์์ ๋์ถ๋ ๊ฒฐ๊ณผ์ด๋ค. ์ค์ ์ ์ฝ์ด ์๋ ์ผ์์ ์ธ ํฐ์น ํ๊ฒฝ์์ ์ฌ์ฉ์๊ฐ ์์ฃผ ๋น ๋ฅด๊ฒ ๋ด๋ฆฌ๋๋ฅด๊ฑฐ๋ ๊ทน๋๋ก ๋๋ฆฌ๊ฒ ๋๋ฅผ ๋ ์๊ฐฑ์ด ์ง๋ ์ฃผํ์์ ๋ฒ ์ด์ค ์ ์ฐ์ฑ์ด ๊ฒฐํฉํ์ฌ ๋ํ๋๋ ์ญํ์ ๋ ธ์ด์ฆ ๋ฐ ์ธ์ง ์๊ณก ๊ฐ๋ฅ์ฑ์ด ์กด์ฌํ๋ค.
๋จ์ผ ๋๊ตฌ(์ฌ์ง) ๋ฐ ์ ์ฒด ์์ญ์ ๊ตญํ์ฑ: ์คํ์ด ์๊ฐ๋ฝ ๋(index fingertip)์ ์์ง ์์ถ ๋์์๋ง ์ง์ค๋์ด ์์ด, ์๋ฐ๋ฅ ์ ์ฒด๋ก ๋ฌผ์ฒด๋ฅผ ์ฅ๊ฑฐ๋(grasping) ์ธก๋ฉด์ผ๋ก ๋ฌธ์ง๋ฅด๋(shear/lateral rubbing) ์ธํฐ๋์
ํ๊ฒฝ, ์ค๋ฆฌ์ฝ ๊ณ์ด ์ธ์ ์ง๋ฌผ(fabric)์ด๋ ๋๋ฌด ๋ฑ ๋ค๋ฅธ ๋๋ฉ์ธ์ ์ฌ์ง ๋ฒ ์ด์ค ํ๊ฒฝ๊น์ง 100% ์ผ๋ฐํํ๊ธฐ์๋ ๋ณ์๊ฐ ์กด์ฌํ๋ค.
ํฅํ๊ณผ์
๋์ ํฐ์น ํ๋กํ(Dynamic Tracing) ๋ฐ ํ ์ ์ด ์๊ณ ๋ฆฌ์ฆ ๊ณ ๋ํ: ์ฌ์ฉ์๊ฐ ๋๋ฅด๋ ์๋์ ๊ถค์ ์ ์ค์๊ฐ ๊ฐ์งํ์ฌ, ํฐ์น ์๋ ๊ฐ๋ณ์ฑ์ ๋ง์ถฐ ์๊ฐฑ์ด(grain)์ ๋ฐ์ ํ์ด๋ฐ๊ณผ ์งํญ์ ๋์ ์ผ๋ก ์ค์๊ฐ ๋ณด์ (velocity-adaptive grain triggering)ํ๋ ์ ์ํ ํ ํฑ ์์ง์ ๊ตฌ์ถํ ๊ณํ์ด๋ค.
3์ฐจ์ ๋ค์ถ ํ ํฑ ๋ ๋๋ง ๋ฐ VR ์ปจํธ๋กค๋ฌ ์ค์ ๋ฐฐํฌ: ์์ง ์๋ ฅ ์ธ์๋ ๋นํ๋ฆผ(torsion), ์ ๋จ๋ ฅ(shear force)๊น์ง ์์ฉํ๋ ๋ค์ถ ๊ฐ๋ณ ์ ์ฐ์ฑ ์ก์ถ์์ดํฐ ์ํคํ ์ฒ๋ฅผ ๊ฐ๋ฐํ๊ณ , ์ด๋ฅผ ์ค์ VR ์๋ฐ์ด๋ฒ ๊ฒ์์ ๋ฐฉ๋ฐฉ ๋ฐ๋ ์งํ ์ง๊ฐ ํํ, ๊ฐ์ ์๋ฃ ์์ ์๋ฎฌ๋ ์ด์ ์ ์ฅ๊ธฐ(organ) ์ด๊ฐ ์ฌํ ๋ฑ์ ํตํฉํ์ฌ ์ฅ๊ธฐ์ ์ ์ ํ๊ฐ๋ฅผ ์งํํ ์์ ์ด๋ค.