Wen Ying, Yeonsu Kim, Adil Rahman, Erzhen Hu, Geehyuk Lee, Seongkook Heo
Redirected Pinch: Efficient and Comfortable Bare-Hand Interaction for 2D Windows in VR
(Abstract) Virtual Reality (VR) offers portable and flexible workspaces. However, enabling efficient and comfortable interactions without external input devices remains challenging. We propose leveraging redirected input to enable comfortable and touch-like interaction for quick and intuitive control. Our design study revealed that while touch interaction performs well with direct input, its performance degrades significantly under input redirection. In contrast, using pinch improves redirected input by providing self-haptic feedback and reducing input dimensionality, thereby compensating for spatial discrepancies. Based on these findings, we introduce Redirected Pinch, a bare-hand interaction technique that combines input redirection with pinch confirmation. It creates a virtual plane at waist height, remapping hand movements on the plane to a vertical window, with pinch gestures used for confirmation. A user study demonstrated that Redirected Pinch achieves a strong balance of accuracy, efficiency, comfort, and sense of agency across fundamental interactions.
(Introduction) Virtual Reality (VR) is emerging as a promising platform for productivity tasks, offering portable, extendable, and distraction-free workspaces that adapt to user needs and contexts [29, 39, 59, 66, 70]. Commercial headsets like Meta Quest 3 and Apple Vision Pro already support desktop-like workflows while enabling large, reconfigurable virtual screens accessible anywhere. Despite the threedimensional nature of VR environments, many productivity applications continue to rely on two-dimensional (2D) interfaces, such as media editing, document review, and web browsing, as 2D layouts remain familiar, efficient, and cognitively lightweight [12, 13, 24, 28, 36, 75]. However, supporting efficient and comfortable interaction with such 2D windows in VR remains a critical challenge.
Common VR interaction methods, such as using handheld controllers, often lack both precision and comfort during prolonged use [71]. External devices like mice [31], keyboards [30, 40, 42], and tablets [12, 13, 28] can provide more accurate control but at the cost of portability, while also requiring physical surfaces for support. With handtracking integrated into commercial headsets, bare-hand interaction offers a natural, lightweight, and always-available alternative [43, 45], capable of enabling more dexterous input than device-constrained methods [48, 58]. However, bare-hand interaction lacks tangibility and stability, making it inefficient for precise input [34, 49], and it induces fatigue when arms are kept elevated, known as the โGorilla Armโ effect [9, 32].
To reduce fatigue and extend reach, researchers have explored input remapping, which manipulates the spatial relationship between real and virtual hands. The Go-Go technique [64] introduced nonlinear reach extension, and subsequent work has amplified small arm movements for large-window interaction [52] or redirected near-body hand motions onto far-away windows for more comfortable input [16]. While effective for ergonomics, input remapping can introduce visualโmotor mismatch, disrupting proprioception and making mid-air touch less reliable [41, 54]. Complementary work improves input redirection using physical surfaces or custom haptic devices [27, 54], but such approaches reduce mobility and are not always available for everyday VR use.
Self-haptic gestures offer an appealing lightweight alternative [25, 33, 61, 86]. Pinch provides tactile confirmation through thumbfinger contact and supports robust gesture recognition [18, 62, 79], leading to wide adoption in commercial systems [6, 55]. However, the performance benefitsofpincharenotconsistent across contexts. Pinch can be slower, more error-prone, and physically demanding in direct mid-air selection [14, 21, 57], yet has shown benefits in indirect interactions such as text entry [26, 41]. We hypothesize that pinch becomes more beneficial in redirected mid-air interaction than in direct interaction for two reasons: (1) its immediate tactile feedback without external devices, compensating for the uncertainty introduced by spatial remapping, and (2) its reduced input dimensionality, as 3D positioning is transformed into 2D motion with a binary pinch confirmation gesture, mitigating accuracy issues caused by depth perception challenges in redirected spaces.
In this paper, we present Redirected Pinch, a novel bare-hand interaction technique that combines pinch with input remapping to enable efficient and comfortable interaction with 2D windows in VR (Figure 1). Redirected Pinch creates a tilted virtual plane at waist height, decoupling the visual workspace from the control space for more ergonomic interaction [4, 16]. Hand movements relative to this plane are remapped in both position and orientation to the application window, while pinches provide self-haptic confirmation. This design reduces fatigue from elevated postures and compensates for spatial inaccuracies caused by remapping. We developed Redirected Pinch through a design study that explored different input mappings (i.e., direct vs. redirected) and confirmation gestures (i.e., touch vs. pinch). Our findings revealed that while touch worked effectively for direct input, its performance degraded significantly under redirected input. In contrast, pinch enhanced redirected input across accuracy, efficiency, and sense of agency, supporting our hypothesis about the value of pinch in remapped interactions.
Finally, weevaluatedtheperformanceofRedirectedPinchagainst commonly used VR interaction techniques, including direct pinch, gaze pinch, and handray pointer, in both simple selection tasks and more complex docking tasks. The results showed that Redirected Pinch provided the best overall balance of comfort, efficiency, and sense of agency in both tasks. In addition, Redirected Pinch was consistently preferred because it required less effort and offered easier control, particularly during prolonged, complex interactions involving continuous and multi-touch input.
(Conclusion) In this work, we investigated mid-air bare-hand interaction techniques for discrete, continuous, and multi-touch inputs on 2D windows in VR. Through design studies comparing input mapping and confirmation gestures, we propose Redirected Pinch, a novel interaction technique that allows users to perform pinch gestures with comfortable postures while enabling efficient interactions by hand redirection and input remapping. We compared Redirected Pinch with three commonly used methods: Direct Pinch, Gaze Pinch, and Handray Pointer. The comparative evaluation showed that Redirected Pinch provides a strong balance of speed, accuracy, agency, and comfort across both selection and docking tasks. While we tested Redirected Pinch in a controlled study, it can be flexible and adapted to more realistic applications. Future work could explore similar techniques in real-world productivity tasks involving windows of varying sizes, distances, orientations, and multi-window setups.
Gorilla Arm Effect
VR๊ณ ๊ธ(HMD) ์ฐฉ์ฉ ํ ๊ฐ์ ๊ณต๊ฐ ๋ด์ ๋ค์์ 2D ์๋์ฐ๋ฅผ ๋์์ผ๋ก์จ ๊ฐ์ ์คํผ์ค ํ๊ฒฝ์ ๊ตฌ์ถํ ์ ์๋ค.
๋ง์ฐ์ค๊ฐ ์กด์ฌํ์ง ์๋ ์ํ์์, ์ ๋ฉ๋ฆฌ ๊ณต์ค์ ๋ ์๋ ํ๋ฉด์ ํฐ์นํ๋ ค๊ณ ํ์ ์ง์์ ์ผ๋ก ๋ป๋ ํ๋์ ์ด๊นจ์ ํ์ ๊ธ๊ฒฉํ ํผ๋ก๋ฅผ ์ ๋ฐํ๊ณ , ์ด๋ฅผ ๊ณ ๋ฆด๋ผ ํ ํจ๊ณผ๋ผ ํ๋ค.
๊ณ ๋ฆด๋ผ ํ ํจ๊ณผ์ ๊ฐ์๋ฅผ ์ํด ์์ ํธ์ํ๊ฒ ํ๋ฆฌ ๋์ด์์ ์์ง์ด๊ณ ํ๋ฉด์ ์์ ์์ง์์ ๋งคํํ๋ ๊ธฐ์ (์
๋ ฅ ๋ฆฌ๋ค์ด๋ ์
, input redirection)์ ๋์
ํ์์ผ๋, ์์ ์ค์ ์์น์ ๋์ด ๋ณด๋ ์์น๊ฐ ๋ฌ๋ผ์ง๋ค ๋ณด๋ ์ ๋ฐํ ํฐ์น ์กฐ์์ ์ด๋ ค์์ด ์กด์ฌํ๋ค.
Redirected Pinch
๊ณ ๋ฆด๋ผ ํ ํจ๊ณผ์ ํด๊ฒฐ์ ์ํด, ์์ด ํธ์ํ ์์น(ํ๋ฆฌ)์์์ ์ํธ์์ฉ๊ณผ ๊ฐ์ฅ ์ง๊ด์ ์ธ ์์ง(์ง๊ธฐ, pinch) ์ ์ค์ฒ๋ฅผ ๊ฒฐํฉํ์๋ค.
์ฌ์ฉ์์ ํ๋ฆฌ ๋์ด์ ๋์ ๋ณด์ด์ง ์๋ ๊ฐ์์ ์ธ ํ๋ฉด์ด ์๋ค๊ณ ๊ฐ์ ํ ๋, ์ฌ์ฉ์๊ฐ ํ๋ฆฌ ๋์ด์์ ์์ ์์ง์ด๋ฉด, ๊ทธ ๊ถค์ ์ด ๋์์ ์์ง ์๋์ฐ ์ฐฝ์ ํฌ์ธํฐ๋ก ๋งคํ๋๋ค. ํ๋ฆฌ ๋์ด์์์ ์ํธ์์ฉ์ ํตํด, ํ์ ๋์ด ๋ค๊ณ ์์ ํ์๊ฐ ์์ผ๋ฏ๋ก ํ์ ํผ๋ก๋๊ฐ ๊ทน์ํ๋๋ค.
์๊ฐ๋ฝ์ผ๋ก ํ๋ฉด์ ๋๋ฅด๋ ๋๋(touch) ๋์ , ์์ง์ ๊ฒ์ง๋ฅผ ๋ง๋ถ๋ชํ๋ ์ง๊ธฐ(pinch) ์ ์ค์ฒ๋ฅผ ํด๋ฆญ ๋ฐ ํ์ธ ์ ํธ๋ก ์ฌ์ฉํ๋ค. ํ์น(pinch)๋ฅผ ํตํด ๋ณธ์ธ์ ์๊ฐ๋ฝ์ด ๋ง๋ฟ๋ ๋ฌผ๋ฆฌ์ ์ด๊ฐ(self-haptic feedback)์ด ์๊ธฐ๊ธฐ ๋๋ฌธ์, ๋๊ณผ ์์ ์์น๊ฐ ๋ค๋ฅด๋๋ผ๋ ํจ์ฌ ์ ํํ๊ณ ์์ ์ ์ธ ์กฐ์์ด ๊ฐ๋ฅํด์ง๋ค.
ํฐ์น(touch)์ ํ์น(pinch) ๋น๊ต
๋์์ ํ๋ฉด์ ์ง์ ๋๋ฅด๋ ์ง์ ํฐ์น๋ ๋น ๋ฅด๊ณ ์ง๊ด์ ์ด์ง๋ง, ์์น๋ฅผ ์๊ณกํ๋ ๋ฆฌ๋ค์ด๋ ์ ํ๊ฒฝ์์๋ ์ ํ๋๊ฐ ๊ธ๊ฒฉํ ํ๋ฝํ๋ค. ๋ฐ๋ฉด, ์์ง์ ๊ฒ์ง๋ฅผ ๋ง๋๋ ํ์น(pinch)๋ ์๊ฐ๋ฝ๋ผ๋ฆฌ ๋ง๋ฟ๋ ์์ฒด ์ด๊ฐ ํผ๋๋ฐฑ(self-haptic feedback)์ ์ฃผ๊ณ ์ ๋ ฅ์ฐจ์์ ์ค์ฌ์ค์ผ๋ก์จ, ์์น๊ฐ ์๊ณก๋ ์ํฉ์์๋ ์ ํ๋๋ฅผ ๋ณด์ํด์ค๋ค.
Gaze pinch(๋์ผ๋ก ๋ณด๊ณ ์ง๊ธฐ), ๋ ์ด์ ํฌ์ธํฐ ๋ฐฉ์ ๋ฑ๊ณผ ๋น๊ต์คํ ํ ๊ฒฐ๊ณผ, redirected pinch๊ฐ ์๋, ์ ํ๋, ํธ์ํจ, sense of agency(์์ด์ ์ ์ธ์, ๋ด๊ฐ ํ๋ฉด์ ์๋ฒฝํ ํต์ ํ๊ณ ์๋ค๋ ๋๋) ๋ฉด์์ ๊ฐ์ฅ ๋ฐ์ด๋ ๊ท ํ์ ๋ณด์๋ค. ๋ํ, ๋๋๊ทธ, ๋ฉํฐํฐ์น ๋ฑ ๋ณต์กํ๊ณ ์ฐ์์ ์ธ ์์ ์ ์ค๋ ํ ๋, ๋ฎ์ ํผ๋ก๋์ ๋์ ํธ์์ฑ์ ์ค์ผ๋ก์จ ์ฌ์ฉ์์ ์ ํธ๋๊ฐ ๋์๋ค.
Adil Rahman, Wen Ying, Md Aashikur Rahman Azim, Michelle Annett, Seongkook Heo
"It Feels Like I am Invited to Communicate": Mediating Ad-Hoc Bystander-VR User Interruptions Through Proactive Proxies
(Abstract) As VRexpands into public spaces, new challenges emerge around spontaneous interactions between bystanders and unfamiliar VR users. While current VR systems often prioritize user awareness of their physical surroundings, they overlook the social dynamics affecting nearby bystanders. We conducted a deception-based study (N=80) examining how interface availability influences bystandersโ comfort, confidence, and hesitation when interrupting VR users. We compared traditional static interruption interfaces (e.g., button on screen) with a proactive proxy that actively approached bystanders upon detecting interruption intent. Static interfaces, due to insufficient cueing, frequently caused bystander discomfort, leading to hesitant physical interruptions or complete communication avoidance. In contrast, the proactive proxy implicitly conveyed social permission, significantly enhancing bystandersโ comfort and confidence. Our findings provide empirical insights into how bystanders assess availability and initiate interruptions with unfamiliar VR users in shared spaces, offering design implications for VR systems that support bystander agency and comfort during these interactions.
(Introduction) Human interruptions in shared spaces are inevitable- whether it is asking a colleague for their input on a shared project [80] or asking a fellow passenger to move [81]. In these moments, gaining someoneโs attention is essential. However, interruptions become more challenging when individuals are deeply immersed in activities that render their primary communication channels unavailable (e.g., listening to music occupies their hearing, reading occupies their vision, etc.). In such cases, secondary communication channels like peripheral vision and ambient awareness serve as subtle cues for initiating interaction.
When individuals are immersed in virtual reality (VR), this dynamic becomes especially complex because their audiovisual immersion effectively limits the secondary use of these primary communication channels [19, 28, 40, 53]. Several alternative modalities have thus been proposed to manage interruptions in VR [50]. Modern VR headsets, such as the Meta Quest and Apple Vision Pro, incorporate open-ear audio technology that enables users to hear verbal interactions. However, this technology can reduce immersion and may not align with all usersโ preferences [50]. Additionally, such technology can be problematic in shared spaces, as sound leakage can disturb nearby individuals and make VR users selfconscious about audio spillover [50, 63]. Users already seek full auditory isolation when completing focused work in public environments by using noise-canceling headphones [46, 55], and as portable VR headsets become productivity tools, extending this need for isolation to the visual domain is a natural progression. Both the Apple Vision Pro and Meta Quest now include dedicated travel modes for use on airplanes and trains [4, 41], explicitly supporting fully immersive use in shared spaces. Moreover, airlines have also begun providing VR headsets to passengers as part of in-flight entertainment services [16], while academic and public libraries increasingly offer VR headset lending programs and dedicated VR spaces [2, 34]. These developments signal that encounters between unfamiliar individuals (i.e., one immersed in VR and the other needing to interrupt them) will become increasingly common. Critically, because bystander awareness features within headsets remain user-controlled, bystanders have no guaranteed means of initiating contact when VR users opt for full immersion. Touch, while potentially the most direct way to interrupt a VR user, is often socially inappropriate in public settings, where interactions among strangers are governed by implicit boundaries [50, 53]. As VRheadsets transition beyond private environments into shared spaces where bystanders may need to interrupt VR users, developing socially appropriate and effective interruption mechanisms becomes increasingly critical [1, 14, 62, 72, 81].
One approach to this challenge has been to enhance VR usersโ awareness of their surroundings. Previous work has explored embedding awareness systems within VR headsets to help users avoid physical accidents and maintain spatial boundaries [22, 40, 42, 49, 51, 53, 58, 73, 84]. Such systems, however, cannot adequately replicate the dynamics of face-to-face interruptions and the subtle social cues that initiators (i.e., bystanders) carefully weigh when determining whether an interruption is appropriate [44]. By shifting control from initiators to VR users, awareness systems invert the natural negotiation process that governs interpersonal interruptions. Beyond immersion trade-offs and the risk of information overload in public spaces, embedded awareness systems may also prematurely redirect a VR userโs attention to bystanders, triggering unwanted face engagements [20]. Although awareness systems can dynamically adjust how they deliver information by withholding alerts until a verbal interaction is attempted [52], such techniques overlook how bystanders assume othersโ unavailability and make decisions not to engage [76, 80]. In crowded public spaces, VR users may even disable these systems to avoid constant notifications about nearby bystanders, especially when they do not anticipate needing to interact with others.
Although limited, prior work has explored bystander-initiated interruption interfaces, such as the HTCโs Knock Knock feature [71] and a physical doorbell peripheral [81], which offer explicit mechanisms for bystanders to attract a VR usersโ attention. However, these interfaces typically assume that bystanders are aware of their existence, limiting their effectiveness in public spaces where interruptions are often spontaneous and involve unacquainted individuals. While interruptions between acquaintances may be more frequent, stranger-to-stranger interactions represent the most challenging case due to the absence of established social rapport and the heightened psychological cost of initiating contact [12]. Solutions effective in such demanding contexts are also likely to generalize to less demanding scenarios. While prior research underscored the importance of observing spontaneous interactions between unfamiliar individuals [50], little is known about how such encounters unfold when unacquainted VR users share physical spaces. Yet, insights into these dynamics are essential for designing VR headsets that integrate more seamlessly into public environments. In this work, we investigate how bystanders naturally navigate spontaneous interactions with unfamiliar VR users in shared environments.
As studying such unplanned encounters at scale is challenging duetotheir situational constraints, we conducted a deception-based study that recreated these constraints in a controlled setting. Unlike prior work that relied on solicited interruptions [17, 23, 50], acquainted participants [50, 53], or anecdotal reports [53], our study placed 80 unsuspecting bystanders in scenarios requiring urgent interaction with an unfamiliar VR user in a shared space. Our study had two interruption interface conditions. The baseline condition, experienced by 40 participants, used a typical desktop VR setup with a static interruption interface (e.g., doorbell peripheral). For the second condition, we hypothesized that explicitly visible, bystander-initiated interfaces could improve interruption experiences. Inspired by public display research [27, 30], we created a robot-based interruption interface, i.e., a proactive proxy that served as a design probe to examine whether explicit interface discoverability impacted bystander experiences. This deliberately extreme intervention was intended to isolate the effect of proactive presentation on bystander comfort and behavior. This condition was experienced by an additional 40 participants.
During our study, we found that, similar to public kiosk research [30], the participants who encountered the baseline condition suffered from the first-click problem [30], where they failed to notice or use the static interface. In contrast, the participants who encountered the proactive proxy reported increased comfort and engagement. While participants preferred mediated interaction over direct interruption, the static interface was largely ignored. These findings underscore the importance of interface visibility and design in shared VR spaces, offering actionable insights for future VR headset development that better supports both users and bystanders.
This research contributes: (1) Empirical evidence that bystanders feel significant discomfort during spontaneous interactions with unfamiliar VR users in shared spaces, with key insights revealing how physical boundaries, perceived safety, and social dynamics influence their interruption behaviors. (2) Identification of key limitations in existing static interruption interfaces, showing how privacy concerns and low situational awareness hinder their discoverability and usability during spontaneous interactions. (3) Design implications for bystander-aware VR interfaces that support comfortable and effective spontaneous interactions, emphasizing the need to explicitly integrate bystander perspectives into VR system design.
(Conclusion) As VRsystems increasingly extend beyond private spaces, understanding and supporting interactions between VR users and bystanders becomes essential. Our research provides empirical evidence that bystanders often feel significant discomfort when interrupting unfamiliar VR users and that traditional static interruption interfaces frequently go unnoticed during spontaneous encounters. Our proactive proxy interface, on the other hand, demonstrated that explicitly discoverable interfaces can significantly improve bystander comfort and interaction quality by signaling social permission to initiate contact. These findings suggest that as VR systems increasingly move into shared and public environments, their design must also consider the needs of nearby bystanders. In particular, incorporating interfaces that clearly communicate availability and guide bystander attention can help trigger interruptions in a comfortable and socially appropriate way.
We advocate for integrating bystander-centric design into VR systems, demonstrating that giving bystanders agency in managing interruptions enhances comfort and engagement. These interfaces can complement existing awareness systems by offering explicit communication channels without compromising immersion. We outline key design considerations for such systems and suggest that future research explore how to effectively combine these approaches across diverse contexts. As VR adoption expands into offices, transit systems, and public venues, this integration will be vital for creating socially attuned VR experiences that respect the needs of all participants.
It Feels Like๊ฐ ์ ํ์ํ ๊น?
๊ณตํญ, ์ ์ํ, ๋์๊ด ๋ก๋น ๋ฑ์ ๊ณต๊ณต์ฅ์์์ ์ด๋ค ์ฌ๋์ด ๋์ ์์ ํ ๊ฐ๋ฆฌ๋ HMD๋ฅผ ์ฐ๊ณ ๊ฐ์ํ์ค์ ๋ชฐ์ ํด์๋ค.
์ฃผ๋ณ์ ์ง๋๊ฐ๋ ํ์ธ(bystander)์, ์ด VR์ฌ์ฉ์์๊ฒ ๊ธํ๊ฒ ๊ธธ์ ๋ฌป๊ฑฐ๋, "์ฌ๊ธฐ์ ๋นํค์ ์ผ ํด์"๋ผ๊ณ ๋ง์ ๊ฑธ์ด์ผ ํ๋ ์ํฉ์ด๋ค.
ํ์ง๋ง, VR์ฌ์ฉ์๊ฐ ๋์ ๊ฐ๋ฆฌ๊ณ ์์ผ๋ ์ธ์ ๋ง์ ๊ฑธ์ด์ผ ํ ์ง ๋์น๊ฐ ๋ณด์ด๊ณ , ๋ชจ๋ฅด๋ ์ฌ๋์ ๋ชธ์ ํญํญ ์น๊ธฐ๋ ๋ฌด์ฒ์ด๋ ๊ป๋๋ฝ๋ค.
It Feels Likeๅ ง ๋ก๋ด๋น์(proactive proxy, ์ฃผ๋์ ํ๋ก์)์ ์ ์ฒด
์๊ธฐ ๋ฌธ์ ํด๊ฒฐ์ ์ํด VR์ฌ์ฉ์ ์์ ๋ฐํด๊ฐ ๋ฌ๋ฆฐ ์์ฃผ ์์ ์ด๋ํ ๋ก๋ด ๋๋ฐ์ด์ค๋ฅผ ์ธ์๋์๋ค.
์ด ๋ก๋ด์ ์๋จ์๋ ์ค๋งํธํฐ ํ๋ฉด ๊ฐ์ ์ธํฐํ์ด์ค(ํ๋ฉด๊ณผ ๋ฒํผ)๊ฐ ์กด์ฌํ๊ณ , ์๋์๋ ๋ฐํด, ๋ชจํฐ, ๋ผ์ฆ๋ฒ ๋ฆฌ ํ์ด๊ฐ ๋ค์ด์๋ค, ref., Appendix C.
"์๋๋ฅผ ๊ฐ์งํด ๋ฅ๋์ ์ผ๋ก ๋ค๊ฐ์จ๋ค"์ ์ค์ ์ ๊ณผ์
์๋๊ฐ์ง: ๋ก๋ด์ด๋ย ์นด๋ฉ๋ผ ๋ฑ์ ์ฃผ๋ณ์ผ์๊ฐ VR์ฌ์ฉ์๊ฐ ์๋, ๊ทธ ์ฃผ๋ณ์ ์์ฑ๊ฑฐ๋ฆฌ๊ฑฐ๋ VR์ฌ์ฉ์๋ฅผ ์ณ๋ค๋ณด๋ฉฐ ๋ค๊ฐ์ค๋ ค๋ ํ์ธ(bystander)๋ฅผ ํฌ์ฐฉํ๋ค, i.e., ๋ง์ ๊ฑธ๊ณ ์ถ์ดํ๋ ์ฌ๋์ ํ๋ํจํด์ ๊ฐ์งํ๋ค.
๋ก๋ด์ด ์ถ๋ฐ: VR์ฌ์ฉ์ ์์ ๊ฐ๋งํ ์ ์๋ ๋ก๋ด์, ํ์ธ์ ์๋๊ฐ ๊ฐ์ง๋๋ฉด(๋ง์ ๊ฑธ๊ณ ์ถ์ดํจ), ํ์ธ ์์ผ๋ก ์ค๋ฅด๋ฅต ์ด๋ํ๋ค.
์ํต์ฃผ์ , ํ๋ฉดํ์: ํ์ธ ์์ ๋์ฐฉํ ๋ก๋ด์ ํ๋ฉด์๋ "์ง๊ธ VR์ฌ์ฉ์๋ ๊ฐ๋ฒผ์ด ๊ฒ์ ์ค์ด๋ผ ๋ง์ ๊ฑฐ์ ๋ ๊ด์ฐฎ์ต๋๋ค. ์ด ๋ฒํผ์ ๋๋ฅด๋ฉด VR์ฌ์ฉ์์๊ฒ ์๋ฆผ์ด ๊ฐ๋๋ค" ๊ฐ์ ์๋ด๋ฅผ ํ์ถํ๋ค.
์ํธ์์ฉ: ํ์ธ์ด ๋ก๋ด์ ํ๋ฉดๅ
ง ๋ฒํผ์ ๋๋ฅด๋ฉด, VR์ฌ์ฉ์์ ๊ณ ๊ธ ํ๋ฉดๅ
ง "์ฃผ๋ณ์ ์๋ ์ฌ๋์ด ๋ํ๋ฅผ ์์ฒญํ์ต๋๋ค"๋ผ๊ณ ์์ ํ๊ฒ ์๋ฆผ์ด ํ์ถ๋๊ณ , VR์ฌ์ฉ์๊ฐ ๊ณ ๊ธ์ ๋ฒ๊ฑฐ๋ ์ธ๋ถ์นด๋ฉ๋ผ๋ฅผ ์ผ์ ํ์ธ๊ณผ ๋์ ๋ง์ถ๋ฉฐ ๋ํ๋ฅผ ์์ํ๋ค.
VR์ด ๊ณต๊ณต์ฅ์๋ก ๋์ฌ ๋ ๋ฐ์ํ๋ ์ํธ์์ฉ๊ณผ ๊ฒฐ๋ก
์ฌ๋๋ค์ ๋ชจ๋ฅด๋ VR์ฌ์ฉ์์๊ฒ ๋ง์ ๊ฑธ ๋ ์๋นํ ์ฌ๋ฆฌ์ ๋ถํธํจ์ ๋๋๋ค.
๋ฒฝ์ ๋ถ์ ๋ฒจ๊ณผ ๊ฐ์ ๊ธฐ์กด์ ์ ์ /์๋์ ์ธํฐํ์ด์ค(static interface)์ ๊ฒฝ์ฐ, ์ฌ๋๋ค์ด ๋๊ธธ๋ ์ฃผ์ง ์๊ณ ๋ฌด์ํ๋ค.
์์คํ ์ด ๋จผ์ ์กด์ฌ๊ฐ์ ๋๋ฌ๋ด๋ฉฐ ๋ค๊ฐ์ค๋ ์ฃผ๋์ ํ๋ก์(proactive proxy) ๋ฐฉ์์ ๊ฒฝ์ฐ, ์ฃผ๋ณ์ธ์๊ฒ "์ง๊ธ ์ํตํด๋ ๋๋ค"๋ ์ฌํ์ ํ๋ฝ(social permission)์ ์ ํธ๋ก ์ค์ผ๋ก์จ, ์ํต์ ์์ด์ ํธ์ํจ์ ์ค๋ค.
Erzhen Hu, Frederik Brudy, David Ledo, George Fitzmaurice, Fraser Anderson
PrevizWhiz: Combining Rough 3D Scenes and 2D Video to Guide Generative Video Previsualization
(Abstract) In pre-production, filmmakers and 3D animation experts must rapidly prototype ideas to explore a filmโs possibilities before fullscale production, yet conventional approaches involve trade-offs in efficiency and expressiveness. Hand-drawn storyboards often lack spatial precision needed for complex cinematography, while 3D previsualization demands expertise and high-quality rigged assets. To address this gap, we present PrevizWhiz, a system that leverages rough 3D scenes in combination with generative image and video models to create stylized video previews. The workflow integrates frame-level image restyling with adjustable resemblance, time-based editing through motion paths or external video inputs, and refinement into high-fidelity video clips. A study with filmmakers demonstrates that our system lowers technical barriers for film-makers, accelerates creative iteration, and effectively bridges the communication gap, while also surfacing challenges of continuity, authorship, and ethical consideration in AI-assisted filmmaking.
(Introduction) Previsualization (previz) is a central practice in filmmaking, enabling directors and creative teams to explore the visual and narrative structure of a scene before production [54]. By creating early visualizations, filmmakers can test ideas for camera angles, blocking, pacing, and emotional beats without the expense of full-scale sets, actors, or detailed assets. Beyond its role as a creative sketching tool, previz also functions as a collaborative artifact, helping directors, cinematographers, production designer, and other stakeholder align around a shared vision [2, 4, 18].
Despite its importance, existing approaches force filmmakers to make trade-offs between speed, fidelity, and control. Storyboards and moodboards are quick and expressive, allowing for early exploration and communication of creative intent [16, 44]. However, these are static, offering limited spatial and temporal representation: they cannot adequately represent motion or timing, making it difficult to visualize complex shots or sequences. 3D previz tools on the other hand allow filmmakers to compose scenes, experiment with camera blocking, and ensure continuity across shots [12, 34, 54]. However, these tools require high-fidelity 3D assets, rigging, and animation expertise [40]. Many existing 3D previz tools also fail to convey fine-grained nuances like emotional beats and microactions.
Recent advances in generative AI can accelerate previsualization by producing images or videos directly from textual prompts, allowing filmmakers to quickly generate outputs with a compelling visual style [17]. Yet, they pose new challenges. Text-to-image and text-tovideo models often struggle with temporal consistency, making coherent motion across frames challenging [36]. They also lack spatial grounding: precise placement of objects and camera, blocking, and continuity are difficult to control. As a result, current approaches using generative AI risk producing highly polished looking results that are disconnected from the filmmakerโs intended structure. Filmmakers need a lightweight and flexible approach that combines the spatial grounding of 3D tools with the expressive richness of generative video tools.
We present PrevizWhiz, a system that allows filmmakers to rapidly explore and visualize their shots by combining rough 3D scene blocking for timing and spatial structure, 2D video references for detailed character motion, and generative stylization guided by images and text. Filmmakers begin by arranging rough 3D proxies to establish prop positions, character movement, as well as camera paths (Figure 1a). They can restyle frames from their 3D scenes to experiment with different aesthetic styles, ranging from strict adherence to loose reinterpretation of their compositions (Figure 1a). Finally, PrevizWhiz allows filmmakers to specify three levels of motion fidelity: (1) coarse motion from 3D blocking, (2) stylized animations that combine motion from 3D blocking with the restyled frame, and (3) control-video animation that augments the stylized animation with 2D reference videos for detailed character motion (Figure 1b). These frames of scene composition, and time-based elements can guide the video generation of final outputs (Figure 1c), shaping style, lighting, composition, and movement in ways that balance the structural consistency of 3D blocking with the expressiveness of 2D generative tools to create previsualization for film.
Our contributions are: (1) PrevizWhiz, a system that combines rough 3D blocking, frame stylization, and granular animation control to enable lightweight yet expressive previsualization, and (2) findings from a user study with filmmakers and 3D artists showing how the system enables rapid ideation during previsualisations and probes their thoughts on generative tools for pre-production.
(Conclusion) We presented PrevizWhiz, a system that combines rough 3D scene blocking, detailed character motion, and video stylization through generative AI to support flexible, rapid previsualization. Through a user study with filmmakers, we found that the system enabled lightweight scene setup, iterative refinement, and expressive authoring across modalities. Our findings suggest that AI-assisted previz can augment creative practice, lowering barriers for independent creators to communicate their creative intent. At the same time, issues of latency, consistency, and fear of displacement highlight the need for careful future design.
Previz
์ํ ๋๋ ์ ๋๋ฉ์ด์ ์ ๋ณธ๊ฒฉ์ ์ผ๋ก ์ดฌ์ํ๊ธฐ ์ ์, ์นด๋ฉ๋ผ ์ต๊ธ, ๋ฐฐ์ฐ์ ๋์ , ํ์ด๋ฐ, ๊ฐ์ ์ ๋ฑ์ ๋ฏธ๋ฆฌ ์๊ฐํํ๋ ํ์์์ ์ด ํ๋ฆฌ๋น์ฆ์ด๋ค. ์ด๋ฅผ ํตํด, ์ ์์ง ๊ฐ์ vision ์ผ์น๊ฐ ๊ฐ๋ฅํด์ง๋ค.
๊ธฐ์กด ํ๋ฆฌ๋น์ฆ ๋ฐฉ์์๋, ์๊ทธ๋ฆผ ์คํ ๋ฆฌ๋ณด๋, 3D ํ๋ฆฌ๋น์ฆ ํ๋ก๊ทธ๋จ, ๊ธฐ์กด ํ ์คํธ ๊ธฐ๋ฐ ์์ฑํAI๊ฐ ์์ผ๋ฉฐ, ๊ฐ๊ฐ์ trade-offs๋ ํ๊ธฐ์ ๋์ผํ๋ค.
์๊ทธ๋ฆผ ์คํ ๋ฆฌ๋ณด๋: ๋น ๋ฅด๊ณ ๊ฐ์ ํํ์ ์์ด์๋ ํจ์จ์ ์ด์ง๋ง, ์ ์งํ๋ฉด์ด๋ผ ์นด๋ฉ๋ผ์ ๋ณต์กํ ์์ง์ ๋๋ ๊ณต๊ฐ์ ์ผ๋ก ์ ํํ ํ์ด๋ฐ ํํ์ ์์ด์ ํ๊ณ๊ฐ ์๋ค.
3D ํ๋ฆฌ๋น์ฆ ํ๋ก๊ทธ๋จ: ๊ณต๊ฐ ๋ฐ ์นด๋ฉ๋ผ ๋์ ์ ์๋ฒฝํ ์ ์ดํ ์ ์๊ธด ํ์ง๋ง, ์ ๊ตํ 3D assets(์บ๋ฆญํฐ, ๋ฐฐ๊ฒฝ ๋ฑ) ๋ฐ rigging(์กฐ์), ์ ๋๋ฉ์ด์ ์ ๋ฌธ๊ฐ๊ฐ ํ์ํ๋ค, i.e., ์ง์ ์ฅ๋ฒฝ์ด ๋๋ฌด ๋๊ณ ์์ ์ ์์ด์ ์๊ฐ์ด ๋ง์ด ์์๋๋ค.
๊ธฐ์กด ํ
์คํธ ๊ธฐ๋ฐ ์์ฑํAI: ๊ทธ๋ด์ธํ๊ณ ๋ฉ์ง ์์์ ๊ธ๋ฐฉ ๋๋ฑ ๋ง๋ค์ง๋ง, ํ๋ ์ ๊ฐ์ ์ผ๊ด์ฑ์ด ๋จ์ด์ง๊ณ , ๊ฐ๋
์ด ์ํ๋ ์ ํํ ์์น์ ์ฌ๋ฌผ์ ๋ฐฐ์นํ๊ฑฐ๋ ์นด๋ฉ๋ผ ์ด๋์ ์ ์ดํ๊ธฐ๊ฐ ๋ถ๊ฐ๋ฅํ๋ค, i.e., grounding์ด ๋ถ์กฑํ๋ค.
Grounding ๋ถ์กฑ
Grounding์ด๋, ์ค์ ๋ฌผ๋ฆฌ์ ๊ท์น์ด๋ ํ์ค์ธ๊ณ์ ๋ฐ์ดํฐ(์ขํ, ๊ตฌ์กฐ, ๋ฌผ๋ฆฌ๋ฒ์น ๋ฑ)์ ๋จ๋จํ ๋ฐ์ ๋ถ์ด๊ณ ๊ณ ์ ํ๋ ๊ฒ์ ๋งํ๋ฉฐ, AI์ grounding ๋ถ์กฑ์ด๋, AI๊ฐ ํ์ค์ ๊ตฌ์ฒด์ ๊ธฐ์ค(์์น, ํฌ๊ธฐ, ๋ฌผ๋ฆฌ๋)์ ๋ฌด์ํ๊ณ , ์๊ธฐ ๋ง์๋๋ก ๊ทธ๋ด์ธํ ์ด๋ฏธ์ง๋ฅผ ์์ํด์ ๋ป์ด๋๊ฐ๋ ๊ฒ์ ๋งํ๋ค.
๋ค์์, ํ ์คํธ ๊ธฐ๋ฐ AI์ grounding ๋ถ์กฑ ์์์ด๋ค: Sora ๋ฑ์ ํ ์คํธ-๋น๋์ค AI์๊ฒ "์นดํํ ์ด๋ธ ์์ ์๋ฉ๋ฆฌ์นด๋ ธ ์์ด ๋์ฌ์๊ณ , ์นด๋ฉ๋ผ๊ฐ ์ค๋ฅธ์ชฝ์ผ๋ก ์ด๋ํ๋ค"๋ผ๊ณ ์ ๋ ฅํ๋ฉด, ์์ ์์ฒด๋ ์์ฃผ ๊ฐ๊ฐ์ ์ด๊ณ ๋ฉ์ง๊ฒ ์์ฑ๋๋ค.
ํ์ง๋ง, ๊ฐ๋ ์ด ์ํ๋ ์ ๋ฐํ ์์ค์ ์ ์ด๊ด์ ์์ ๋ณด๋ฉด, ๋ค์๊ณผ ๊ฐ์ ๋ฌธ์ ์ ์ด ์กด์ฌํ๋ค
์์น์ ์ด ๋ถ๊ฐ: "์๋ฉ๋ฆฌ์นด๋ ธ ์์ ํ ์ด๋ธ ์ ์ค์์์ ์ ํํ ์ผ์ชฝ์ผ๋ก 15cm ๋จ์ด์ง ๊ณณ์ ๋ฐฐ์นํด ์ค"๋ผ๊ณ ํด๋, AI๋ "์ผ์ชฝ 15cm"๋ผ๋ ๋ฌผ๋ฆฌ์ ์ขํ๋ฅผ ์ดํดํ์ง ๋ชปํด ์๋ฑํ ๊ณณ์ ์์ ๋ฐฐ์นํ๋ค.
๋ฌผ๋ฆฌ์ ์ผ๊ด์ฑ ๋ถ๊ดด: ์นด๋ฉ๋ผ๊ฐ ์ค๋ฅธ์ชฝ์ผ๋ก ์ด๋ํ๋ ๋์, ํ ์ด๋ธ ๋ค์ ์๋ ์์ฅ์ ํฌ๊ธฐ๊ฐ ๊ฐ์๊ธฐ ์ปค์ง๊ฑฐ๋, ์์ ์์ก์ด ๋ฐฉํฅ์ด ์๋ฑํ๊ฒ ๋ฐ๋๋ ๋ฑ ํ๋ ์ ๊ฐ์ ์ผ๊ด์ฑ์ด ๋ถ๊ดด๋๋ค.
I.e., ํ
์คํธ๋ผ๋ ์ถ์์ ๋ช
๋ น๋ง์ผ๋ก๋ ํ์ค์ 3์ฐจ์๊ณต๊ฐ ๋ฒ์น๊ณผ ์ ํํ ์์น์ ์์น๋ฅผ AI์๊ฒ ๊ณ ์ (grounding)์ํฌ ์ ์๋ค.
PrevizWhiz
PrevizWhiz๋, ๋์ถฉ ๋ฐฐ์นํ 3D๋ชจํ(๊ณต๊ฐ ๊ฐ์ด๋), ์ผ๋ฐ 2D๋น๋์ค(๋์ ๊ฐ์ด๋), ์์ฑํAI(์คํ์ผ ์ ํ๊ธฐ)๋ฅผ ๊ฒฐํฉํ๋ค.
Rough 3D Blocking: ์ ๊ตํ assets์ด ์๋, ๋ค๋ชจ/์ธ๋ชจ ๋ฑ์ ๊ฑฐ์น ํ๋ก์(๋ชจํ)๋ค๋ก ์์น์ ์นด๋ฉ๋ผ ๋์ , ํ์ด๋ฐ๋ง ๋์ถฉ ์ก์๋๋ค, i.e., ๊ณต๊ฐ์ ๋ผ๋๋ฅผ ๊ตฌ์ถํ๋ค.
2D๋น๋์ค ๋ ํผ๋ฐ์ค ์ตํฉ: ์บ๋ฆญํฐ์ ์ธ๋ฐํ ์์ง์์ด๋ ๊ฐ์ ๋ฌ์ฌ๋ ์ผ๋ฐ 2D์นด๋ฉ๋ผ๋ก ๋์ถฉ ์ฐ์ ๋น๋์ค๋ฅผ ์์ค๋ก ํ์ฉํ๋ค, i.e., ๊ฐ๋น์ผ 3D ์ ๋๋ฉ์ด์ ๊ณผ์ ์ด ์๋ต๋๋ค.
์์ฑํAI ์คํ์ผ๋ง ์ง์: AI๋ชจ๋ธ์ด 3D๋ผ๋์ 2D๋น๋์ค๋ฅผ ๊ฐ์ด๋๋ผ์ธ์ผ๋ก ํ์ฉํ์ฌ, ๊ฐ๋
์ด ์ํ๋ ๋ฉ์ง ํํ(์ค์ฌํ, ์ ๋๋ฉ์ด์
ํ ๋ฑ)์ ๊ณ ํ์ง/๊ณ ํ์ง ๋น๋์ค ํด๋ฆฝ์ผ๋ก ๋ ๋๋งํ๋ค, i.e., 3D์ ๊ณต๊ฐ ํต์ ๋ ฅ๊ณผ ์์ฑํAI์ ํํ๋ ฅ/์๋ ๋ฉด์์ ์ฅ์ ์ ํ๋ณดํ ์ ์๋ค.
์ค์ ์ํ๊ฐ๋ ๋ฐ 3D์ํฐ์คํธ๋ค์ ๋์์ผ๋ก ์ ์ ์คํฐ๋๋ฅผ ์งํํ ๊ฒฐ๊ณผ
๊ธฐ์ ์ ์ฅ๋ฒฝ์ ๋ฎ์ถ๋ฉด์, ์์ด๋์ด๋ฅผ ๋น ๋ฅด๊ฒ ์๋ํด๋ณด๋ ์ฐฝ์์ ๋ฐ๋ณต ์์ (creative iteration)์ ๊ฐ์ํ ํ ์ ์์๋ค. ๋ํ, ๋ ๋ฆฝ์ํ์ ์์๋ค์ ๊ฒฝ์ฐ, ํฐ ๋น์ฉ์ ๋ค์ด์ง ์์ผ๋ฉด์๋ ์์ ์ ๋น์ ์ ์๊ฐํํ ์ ์์๋ค.
๋ค๋ง, ์์ฑํAI์ ๊ณ ์ง์ ๋ฌธ์ ์ธ ์ฐ์์ฑ(continuity), ์ ์๊ถ(authorship), ์ค๋ฆฌ์ ๊ณ ๋ ค์ฌํญ(ethical consideration, ์ธ๋ ฅ๋์ฒด ์ฐ๋ ค ๋ฑ)์ด๋ผ๋ ๊ณผ์ ๋ค์ ํจ๊ป ๋์ถํ๋ค.
Adil Rahman, Koichiro Ninuma, Aakar Gupta
DataSpeck: An AI-Driven Human-in-the-Loop System for Automating Transformations in Data Conversion Workflows
(Abstract) In data-driven systems, integrating disparate data sources becomes challenging when incoming data does not conform to the systemโs specifications. Despite advances in automated schema matching systems, data integration tasks involving complex semantic interrelationships still require users to manually identify and define transformations between datasets, which can be cognitively demanding and time-consuming. We present DataSpeck, an end-to-end system that automates the conversion of disparate data sources to fit any pre-existing data specification. DataSpeck employs an AI-driven human-in-the-loop design, using LLMs to analyze semantic relationships and generate step-by-step transformation pipelines autonomously, while only requesting user attention to resolve semantic ambiguities. In our technical evaluation, DataSpeck successfully automated ~86% of varied data transformations while generating interpretable strategies with confidence scores and targeted clarification requests. In a user study (N=12), participants completed data conversion tasks ~53% faster with significantly reduced cognitive load using DataSpeck compared to Microsoft Excel with Copilot.
(Introduction) Data-driven systems are often built around specific data formats and structures optimized for their intended use cases. However, the growing variety of data sources poses a major challenge, as incoming data frequently deviates from expected formats and must be adapted before it can be utilized [58]. For example, a healthcare analytics platform designed to process standardized patient records may struggle to integrate external research datasets with differing schemas [11]. Integrating new data sources within established data architectures requires manually designing elaborate transformation pipelines which map incoming data to match the specification format [113]. This process consumes time and effort and must be repeated every time a new data source is introduced - creating significant bottlenecks in data integration workflows.
Reconciling different data formats has been a longstanding research goal spanning decades in data management and information systems, with schema matching and mapping techniques being the primary approaches to enable interoperability between heterogeneous data formats [85, 98]. While these systems excel at finding match candidates between source and target schemas, they often leave the semantic relationships and necessary transformations between matched attributes for users to define manually [6, 10, 13, 98, 126]. In data workflows, this transformation step remains a significant bottleneck, requiring detailed preparation before data can be utilized [75]. To address this, various interaction techniques seek to simplify the process by reducing manual coding requirements. Programming-by-example [8, 20, 44, 49, 62, 106] and Programming-by-demonstration [47, 64] systems have proven particularly effective in eliminating the need to manually write transformation code and allow users to automate data formatting through demonstrated examples. Natural language interfaces [56] have further enabled users to describe transformation requirements in plain language, though they often require precise phrasing to avoid ambiguity [66, 109, 123]. Despite these advancements, users must still understand data structures, interpret relationships, and devise appropriate transformationsโa process that remains cognitively demanding for complex or unfamiliar datasets. While human insight remains crucial for exploratory analysis [105], scenarios with predetermined data structures could benefit from more nuanced automation approaches that reduce unnecessary overhead when adapting diverse sources to existing specifications.
In this paper, we present DataSpeck, an AI-driven system designed towards converting a dataset into a prescribed specification. Unlike previous approaches, DataSpeck does not require users to provide examples or describe the transformation process explicitly. Instead, it leverages LLMs to analyze the semantic relationships between the new data source and the pre-existing data specification, and uses this understanding to automatically generate both transformation strategies and the corresponding data conversion scripts. However, data conversions can involve ambiguities that require additional context. To address such scenarios, DataSpeck incorporates a human-in-the-loop design, classifying the need for human input based on system confidence, and prompting users to provide additional context when necessary.
To evaluate the effectiveness of our human-in-the-loop system, we first conducted a technical evaluation to understand the boundaries of our systemโs automation capabilities. We tested DataSpeck against 43 isolated transformation scenarios and 5 complex realworld data conversion scenarios. Our system was able to successfully automate ~86% of the data transformation operations and yielded appropriate system confidence scores and clarification requests. Our technical evaluation also highlighted transformation scenarios where the automation capabilities struggled. Using these insights, we designed a user study with 12 participants who had prior data integration experience to measure the systemโs impact on user performance and effort. As a baseline, we compared DataSpeck against Microsoft Excel with Copilot, which represented the counterpart human-driven, AI-in-the-loop approach where users infer the semantic relationships on their own, and then use natural language interactions to design the transformations. Participants achieved significantly higher performance and efficiency in the given data conversion tasks, completing tasks 53% faster on average using DataSpeck, and reported significantly lower levels of mental demand, effort, and frustration on the NASA-TLX scale. They found DataSpeckโs ability to automatically infer transformations from source-specification pairs highly usable and valuable for their professional settings, noting that such an interaction could transform the tedious task of manually analyzing datasets and writing migration scripts into simply reviewing system-generated strategies and answering clarifications, thereby significantly reducing manual effort and cognitive load.
We summarize our contributions as follows: (1) The design of DataSpeck, an end-to-end system that automates the transformation of disparate data sources into preexisting data specifications through an AI-driven, human-inthe-loop approach. (2) Technical findings demonstrating that DataSpeck successfully automates a comprehensive set of data transformation operations across diverse scenarios, while generating appropriate confidence scores and targeted clarification requests when human input is needed. (3) User performance results showing that DataSpeckโs AI-driven, human-in-the-loop approach enables users to complete data conversion tasks more efficiently while significantly reducing cognitive load across all NASA-TLX dimensions compared to human-driven AI-in-the-loop approaches.
(Conclusion) Data conversions often require significant manual effort to align diverse data sources to specific structural requirements. Traditional methods do not directly solve for this problem and using existing tools requires extensive manual efforts; in contrast, DataSpeck employs an end-to end AI-driven human-in-the-loop approach that understands the data, strategizes the conversion pipeline, and executes the transformation steps in a highly transparent manner, employing a tiered confidence-based human intervention mechanism when it may be needed. A technical evaluation shows that DataSpeck was able to handle a large number of transformation tasks without human intervention, automate a large part of the pipeline for realworld end-to-end conversions, and successfully surface the need for human intervention for the rest. Our user study demonstrates that DataSpeckโs AI-driven human-in-the-loop approach is significantly more efficient in terms of time and effort compared to a familiar human-driven AI-assisted method. While fully human-driven methods offer complete control, they are often time-consuming and cognitively demanding. In contrast, a human-in-the-loop approach, where the AI automates the majority of transformations and requests user input for clarification or uncertainties, demonstrated potential to significantly enhance efficiency and reduce cognitive overheads for such data tasks.
๋ฐ์ดํฐ ์ค์ฌ ์์คํ ๊ณผ ํฌ๋งท/๊ตฌ์กฐ(specification)
๋ฐ์ดํฐ ์ค์ฌ ์์คํ ๋ค์ ์์ ๋ค์ ๋ชฉ์ ์ ๋ง๊ฒ ์ต์ ํ๋ ๊ณ ์ ์ ํฌ๋งท/๊ท๊ฒฉ(Specification)์ ๊ฐ์ง๊ณ ์๋ค.
ํ์ง๋ง, ์ฌ๋ฌ ์ธ๋ถ ์์ค์์ ์๋ก์ด ๋ฐ์ดํฐ๋ค์ด ๋ค์ด์ฌ ๋, ๊ธฐ์กด ๊ท๊ฒฉ๊ณผ ์ผ์นํ์ง ์๋ ๊ตฌ์กฐ์ /์๋ฏธ๋ก ์ ๋ถ์ผ์น ๋ฌธ์ ๊ฐ ์์ ๋ฐ์ํ๋ค.
์ด๋ฅผ ํด๊ฒฐํ๊ธฐ ์ํด ๊ธฐ์กด ์์คํ
์ ๋ง๊ฒ ๋ฐ์ดํฐ๋ฅผ ๋งคํ/๋ณํ(data conversion)ํ๋ ํ์ดํ๋ผ์ธ์ ์๋์ผ๋ก ์ค๊ณํด์ผ ํ๋๋ฐ, ์ด ๊ณผ์ ์ ์๋นํ ์๊ฐ ๋ฐ ๋
ธ๋ ฅ์ ์๋ชจํ๋ค, i.e., ๋ฐ์ดํฐ ํตํฉ ์ํฌํ๋ก์ฐ์ ๋ณ๋ชฉ ์์ธ์ด๋ค.
๊ธฐ์กด ๊ธฐ์ ๋ค์ trade-offs
์๋ ์คํค๋ง ๋งค์นญ ์์คํ : source schema์ target schema์ ํ๋(์ปฌ๋ผ)๊ฐ ์ฐ๊ฒฐ(match candidates)ํ๋ ๊ฒ์ ์ ํ์ง๋ง, ๊ตฌ์ฒด์ ์ผ๋ก ๋ฐ์ดํฐ๋ฅผ ์ด๋ป๊ฒ ํตํฉ, ๋ถํ , ๋จ์๋ณํํด์ผ ํ๋์ง ๋ฑ์ ์๋ฏธ๋ก ์ ๊ด๊ณ(semantic interrelationships)์ ์ค์ ๋ณํ๋ก์ง์ ๋ํด์๋, ์ฌ์ฉ์์ ์์์ (์ฝ๋ฉ ๋ฐ ์ ์)์ด ํ์ํ๋ค.
์์ ๊ธฐ๋ฐ ํ๋ก๊ทธ๋๋ฐ(PBE, Programming By Example) ๋ฐ ์์ฐ ๊ธฐ๋ฐ ์์คํ : ์์๋ฅผ ๋ณด์ฌ์ฃผ๋ฉด ์ฝ๋๋ฅผ ์๋์ผ๋ก ๋ง๋ค์ด์ฃผ์ด ์์์ (์ฝ๋ฉ)์ ์ค์์ง๋ง, ์ฌ์ฉ์๊ฐ ๋ฐ์ดํฐ ๊ตฌ์กฐ๋ฅผ ์๋ฒฝํ ์ดํดํ ์ํ์์ ์ง์ ์ ํํ ์ ์ถ๋ ฅ ์์๋ฅผ ๋ง๋ค์ด ์ ๊ณตํด์ผ ํ๋ฏ๋ก, ์ ์ ์ ๋ถ๋ด(cognitive demand)์ ์ฌ์ ํ ์กด์ฌํ๋ค.
์์ฐ์ด ์ธํฐํ์ด์ค: ๋ง๋ก ๋ช
๋ นํ๋ฉด ๋ฐ์ดํฐ ๋ณํ์ ์ํํ์ง๋ง, ์ฌ์ฉ์๊ฐ ๊ทน๋๋ก ์ ๋ฐํ๊ณ ์ ํํ ์์ฐ์ด ๋ฌธ์ฅ(prompt engineering)์ ๊ตฌ์ฌํด ์ง์ํด์ผ ํ๋ฏ๋ก, ๋ฐ์ดํฐ์
์ด ๋ณต์กํ ๊ฒฝ์ฐ ํจ์จ์ฑ์ด ๋จ์ด์ง๋ค.
๋ฐ์ดํฐ ๋ณํ(data conversion)์์์ semantic gap ๋ถ์กฑ
Semantic gap์ด๋, ๋ฐ์ดํฐ์ ํ๋๋ช (์, buyer, customer)์ด๋ ๋ฐ์ดํฐ์ ํํ(์, ์ ์ฒด์ฃผ์ ๋ฌธ์์ด, ์ฐํธ๋ฒํธ/๋์ ๋ถํ ํ๋)๊ฐ ์๋ก ์์ดํ ๋, ์ด๊ฒ๋ค์ด ๋ณธ์ง์ ์ผ๋ก๋ ๋์ผํ ์๋ฏธ๋ฅผ ์ง๋๊ณ ์์์ ํ์ ํ๊ณ ์ฐ๊ฒฐํด์ฃผ๋ ๋ ผ๋ฆฌ์ ๋ฐํ์ ์๋ฏธํ๋ค.
๊ธฐ์กด AI์ ์๋ฏธ๋ก ์ ์ดํด ๋ถ์กฑ: ์ผ๋ฐ์ ์ธ AI ์ด์์คํดํธ๋ ๋จ์ผ ํ์ ์์์ ๋ง๋ค๊ฑฐ๋ ๊ฐ๋จํ ํฌ๋งทํ
์ ๋๋ ๋ฐ์๋ ๋ฐ์ด๋๋ค. ํ์ง๋ง, ์ฌ๋ฌ ํ
์ด๋ธ์ ํฉ์ด์ง ๋ฐ์ดํฐ(์, ์ ํ๋ฌด๊ฒ๋ก ํฌ์ฅ์กฐ๊ฑด ๋ถ๋ฅ, ํ์จ ์ ์ฉ ๋ฑ) ๊ฐ์ ๊ณ ์ฐจ์์ ์ธ๊ณผ๊ด๊ณ ๋ฐ ๋งคํ์ ๋ต์ ์์จ์ ์ผ๋ก ์์ฑํ์ง๋ ๋ชปํ๋ค, i.e., ์ฌ์ฉ์๊ฐ ์ฒ์๋ถํฐ ๋๊น์ง ๋ณํ์ง์๋ฅผ ๋ฆฌ๋ํด์ผ ํ๋ฏ๋ก, ์์ํ๊ฑฐ๋ ๊ฑฐ๋ํ ๋ฐ์ดํฐ์
์ ๋ง์ฃผํ๋ฉด, AI๊ฐ ์์ด๋ ์์
์ด ์ค๋จ๋๋ค.
DataSpeck ๋ฐ human-in-the-loop
DataSpeck์ ์ฌ์ฉ์๊ฐ ์ง์ ์์๋ฅผ ์ฃผ๊ฑฐ๋ ์ผ์ผ์ด ์ง์๋ฅผ ๋ด๋ฆด ํ์ ์์ด, ์์ค ๋ฐ์ดํฐ์ ๊ณผ ํ๊ฒ ๋ช ์ธ(specification) ์์ ์ ๋ ฅํ๋ฉด, AI๊ฐ ์ ์ ์ ์ผ๋ก ๋ฐ์ดํฐ ๋ณํ ์คํฌ๋ฆฝํธ ๋ฐ ํ์ดํ๋ผ์ธ์ ์์จ์ ์ผ๋ก ์์ฑํ๋ค. ๋ค๋ง, ์๋ฏธ๊ฐ ๋ชจํธํ failure point์์๋ง ์ธ๊ฐ์๊ฒ ์ง๋ฌธ์ ๋์ง๋, AI์ฃผ๋ํ ์ธ๊ฐ์ฐธ์ฌํ(AI-driven human-in-the-loop) ๋์์ธ์ ์ฑํํ๋ค.
Analysing dataset descriptors (๋ฐ์ดํฐ ์์ฝ ๋ถ์): ๊ฑฐ๋ํ ๋ฐ์ดํฐ๋ฅผ LLM์ ๋ค ๋ฃ์ ์๋ ์์ผ๋ฏ๋ก, ๋ฐ์ดํฐ์ ๊ตฌ์กฐ์ ํน์ง(์ํ, ๊ฒฐ์ธก์น, ํต๊ณ๋)์ ์ถ์ถํ ํ, LLM์ ํตํด ๊ฐ ์ปฌ๋ผ์ ์๋ฏธ, ๋ฐ์ดํฐ ํ์ , ๋จ์, ํฌ๋งท์ ๋ด์ ์์ถ๋ semantic descriptors๋ฅผ ์์จ์ ์ผ๋ก ์์ฑํ๋ค.
Establishing Semantic Relationships (์๋ฏธ๋ก ์ ๊ด๊ณ ๊ตฌ์ถ): ์์ค์ ํ๊ฒ ๋ช ์ธ์ descriptors๋ฅผ ๋น๊ต ๋ถ์ํ์ฌ, ํ๊ฒ ๋ช ์ธ์ ๊ฐ ์ปฌ๋ผ์ ์์ค ๋ฐ์ดํฐ๋ก๋ถํฐ ์ด๋ป๊ฒ ์ ๋ํด๋ผ ์ ์์์ง, ์์ฐ์ด๋ก ๋ ๋ณํ์ ๋ต(transformation strategy) ์ธํธ๋ฅผ ์ค์ค๋ก ์๋ฆฝํ๋ค.
Tiered confidence-based human intervention (์ ๋ขฐ๋ ๊ธฐ๋ฐ ์ธ๊ฐ ๊ฐ์ ): AI๊ฐ ์์จ์ ์ผ๋ก ์ ๋ต์ ์ง๋ ๊ณผ์ ์์ ํ๋จ์ ๋ชจํธํจ(ambiguities)์ด ์๊ธฐ๋ฉด ์ด๋ฅผ ์ ๋ขฐ๋์ ๋ฐ๋ผ confident(ํ์ ), assuming(๊ฐ์ ), insufficient(๋ถ์กฑ)์ 3๋จ๊ณ๋ก ๋ถ๋ฅํ๋ค. Assuming์ผ ๋๋ ๊ฐ์ ์ ์๋ฆฝํ ๋ค ํ์ธ์ ์์ฒญํ๊ณ , insufficient์ผ ๋๋ง ์ฌ์ฉ์์๊ฒ ๋ช ํํ ์์ฒญ(clarification request)์ ๋ณด๋ด ๋ชจํธํจ์ ํด๊ฒฐํ๋ค.
Dynamic step/hybrid code generation (๋์ ์ฝ๋ ์์ฑ): ์๋ฆฝ๋ ์ ๋ต์ ๋ฐํ์ผ๋ก ํ ๋จ๊ณ์ฉ ์คํ ๊ฐ๋ฅํ ํ์ดํ๋ผ์ธ Python ์ฝ๋๋ฅผ ๋น๋ํ๋ฉฐ, ์ ๋ฐํ ๊ท์น ํํ์ด ์ด๋ ค์ธ ๋๋ ํ ๋จ์๋ก LLM์ ํธ์ถํ๋ ํจ์(ask_llm)๋ฅผ ์ฝ์
ํ๋ ํ์ด๋ธ๋ฆฌ๋ ์ ๋ต์ ์ทจํด ํ์ดํ๋ผ์ธ์ ์์ฑํ๋ค.
์ค์ ๊ธฐ์ ์ ํ๊ฐ ๋ฐ ์ ์ ์คํฐ๋ ์งํ ๊ฒฐ๊ณผ
๊ธฐ์ ์ ์ฑ๋ฅ ๊ฒ์ฆ: 43๊ฐ์ ๊ฒฉ๋ฆฌ๋ ๋ณํ ์๋๋ฆฌ์ค์ 5๊ฐ์ ๋ณต์กํ ์ค์ Kaggle ๋ฐ์ดํฐ์ ์์ ๋์์ผ๋ก ์๋ํ ์ฑ๋ฅ์ ํ ์คํธํ ๊ฒฐ๊ณผ, ์ธ๊ฐ์ ๊ฐ์ ์์ด ์ ์ฒด ๋ฐ์ดํฐ ๋ณํ ์์ ์ ์ฝ 86%~90%๋ฅผ ์์ ํ ์๋์ผ๋ก ์ฑ๊ณต์์ผฐ์ผ๋ฉฐ, ์ ๋ขฐ๋ ๊ธฐ๋ฐ ์ง๋ฌธ ๋ฉ์ปค๋์ฆ๋ ๋งค์ฐ ์ ํํ๊ฒ ์๋ํจ์ ์ฆ๋ช ํ๋ค. ๋จ, ๋ฐ์ดํฐ ํํฐ๋ง์ด๋ ์ ๋ ฌ ๊ฐ์ด ๊ฒ์ผ๋ก ๋๋ฌ๋์ง ์๋ ๊ท์น ํจํด์ ์ค์ค๋ก ์์์ฑ์ง ๋ชปํ๊ณ ๋์น๋ ํ๊ณ๋ ๋ฐ๊ฒฌ๋์๋ค.
์ ์ ์คํฐ๋ ๊ฒฐ๊ณผ (vs. MS Excel with Copilot): ์ฌ์ฉ์๊ฐ ์ง์ ๋ฐ์ดํฐ ๊ด๊ณ๋ฅผ ํ์ ํ๊ณ AI์๊ฒ ๋ช ๋ น์ ๋ด๋ ค์ผ ํ๋ ์ธ๊ฐ ์ฃผ๋ํ AI ๋ณด์กฐ ๋ฐฉ์(Excel + Copilot)์ ๋์กฐ๊ตฐ์ผ๋ก ์คํ์ ์งํํ๋ค.
์ํ ํจ์จ์ฑ: DataSpeck์ ์ฌ์ฉํ ์ฐธ๊ฐ์๋ค์ด ์์ ์ ํ๊ท 53% ๋ ๋น ๋ฅด๊ฒ ์๋ฃํ๋ค.
์ง๋ฌด ๋ถํ ๊ฐ์: NASA-TLX(์ง๋ฌด๋ถํ ํ๊ฐ์งํ) ๊ธฐ์ค ์ ์ ์ ์๊ตฌ๋, ๋
ธ๋ ฅ, ์ข์ ๊ฐ ๋ฑ์ ๋ชจ๋ ์ฐจ์์์ ์ธ์ง์ ๋ถํ๊ฐ ์ ์๋ฏธํ๊ฒ ๊ฐ์ํ๋ค. ๋ฐ์ดํฐ์
์ ์ง์ ๋ถ์ํ๊ณ ๋ง์ด๊ทธ๋ ์ด์
์คํฌ๋ฆฝํธ๋ฅผ ๋ง๋๋ ๋์ , AI๊ฐ ๋ค ์ง๋์ ์ ๋ต์ ๊ฒํ (reviewer)ํ๊ณ ๋ชจํธํ ์ง๋ฌธ์ ๋ต๋ง ํด์ฃผ๋ ๋ฐฉ์์ผ๋ก ์ํฌํ๋ก์ฐ๊ฐ ํ์ ๋จ์ ํ์ธํ๋ค.
NASA-TLX (NASA Task Load Index, ๋์ฌ ์์ ๋ถํ ํ๊ฐ์งํ)
์ฌ๋์ด ์ด๋ค ์์ (task)์ ์ํํ ๋ ์ ์ ์ ์ผ๋ก ์ผ๋ง๋ ํ๋ค์๋์ง(์ง๋ฌด ๋ถํ, mental workload) ์ธก์ ์ ์ํด NASA์์ ๊ฐ๋ฐํ ์ธ๊ณ ํ์ค ์ค๋ฌธ ํ๊ฐ ๋๊ตฌ์ด๋ค.
๋จ์ํ "์ด ์์คํ ์ฐ๋๊น ํธํด์?"๋ผ๊ณ ๋ญ๋ฑ๊ทธ๋ ค ๋ฌป์ง ์๊ณ , ์ธ๊ฐ์ ์ธ์ง์ ๊ณ ํต๊ณผ ์คํธ๋ ์ค๋ฅผ 6๊ฐ์ง ๊ตฌ์ฒด์ ์ธ ์ฐจ์์ผ๋ก ๋๋์ด ์ ์(0~100์ )๋ฅผ ์ธก์ ํ๋ค.
์ ์ ์ ์๊ตฌ (Mental Demand): ๋จธ๋ฆฌ๋ฅผ ์ผ๋ง๋ ๋ง์ด ์จ์ผ ํ๋๊ฐ? (์๊ฐ, ๊ณ์ฐ, ๊ธฐ์ต ๋ฑ)
์ ์ฒด์ ์๊ตฌ (Physical Demand): ๋ชธ์ ์ผ๋ง๋ ๋ง์ด ์์ง์ฌ์ผ ํ๋๊ฐ? (๋ฐ๊ณ ๋น๊ธฐ๊ธฐ, ํ์น ์กฐ์, ํ์ดํ ๋ฑ)
์๊ฐ์ ์๊ตฌ (Temporal Demand): ์๊ฐ์ ์๋ฐ์ด๋ ์ด๋ฐํจ์ ์ผ๋ง๋ ๋๊ผ๋๊ฐ?
์์ ์ฑ๋ฅ (Performance): ๋ด๊ฐ ๋ชฉํ๋ฅผ ์ผ๋ง๋ ์ฑ๊ณต์ ์ผ๋ก ๋ฌ์ฑํ๋ค๊ณ ๋๋ผ๋๊ฐ?
๋ ธ๋ ฅ (Effort): ์ํ๋ ์ฑ๊ณผ๋ฅผ ๋ด๊ธฐ ์ํด ์ ์ ์ ยท์ ์ฒด์ ์ผ๋ก ์ผ๋ง๋ ์ ๋ฅผ ์จ์ผ ํ๋๊ฐ?
์ข์ ๊ฐ (Frustration): ์์ ์ ํ๋ ๋์ ์ผ๋ง๋ ์ง์ฆ, ์คํธ๋ ์ค, ๋๋ด์ ๋๊ผ๋๊ฐ?