<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Publications &#8211; Immersive Computing Lab</title>
	<atom:link href="https://www.immersivecomputinglab.org/category/publications/feed/" rel="self" type="application/rss+xml" />
	<link>https://www.immersivecomputinglab.org</link>
	<description></description>
	<lastBuildDate>Wed, 08 Jul 2026 01:47:07 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.0.2</generator>
	<item>
		<title>Infinite Gaze Generation for Videos with Autoregressive Diffusion</title>
		<link>https://www.immersivecomputinglab.org/publication/infinite-gaze/</link>
		
		<dc:creator><![CDATA[Jenna Kang]]></dc:creator>
		<pubDate>Sun, 05 Jul 2026 06:37:04 +0000</pubDate>
				<category><![CDATA[Publications]]></category>
		<guid isPermaLink="false">https://www.immersivecomputinglab.org/?post_type=publication&#038;p=3797</guid>

					<description><![CDATA[Predicting human gaze in video is fundamental to advancing scene understanding and multimodal interaction. While traditional saliency maps provide spatial probability distributions and scanpaths offer ordered fixations, both abstractions often collapse the fine-grained temporal dynamics of raw gaze. Furthermore, existing models are typically constrained to short-term windows (3-5s), failing to capture the long-range behavioral dependencies inherent in real-world content. We present a generative framework for infinite-horizon raw gaze prediction in videos of arbitrary length. By leveraging an autoregressive diffusion model, we synthesize gaze trajectories characterized by continuous spatial coordinates and high-resolution timestamps. Our model is conditioned on a saliency-aware visual latent space. Quantitative and qualitative evaluations demonstrate that our approach significantly outperforms existing approaches in long-range spatio-temporal accuracy and trajectory realism.]]></description>
										<content:encoded><![CDATA[Predicting human gaze in video is fundamental to advancing scene understanding and multimodal interaction. While traditional saliency maps provide spatial probability distributions and scanpaths offer ordered fixations, both abstractions often collapse the fine-grained temporal dynamics of raw gaze. Furthermore, existing models are typically constrained to short-term windows (3-5s), failing to capture the long-range behavioral dependencies inherent in real-world content. We present a generative framework for infinite-horizon raw gaze prediction in videos of arbitrary length. By leveraging an autoregressive diffusion model, we synthesize gaze trajectories characterized by continuous spatial coordinates and high-resolution timestamps. Our model is conditioned on a saliency-aware visual latent space. Quantitative and qualitative evaluations demonstrate that our approach significantly outperforms existing approaches in long-range spatio-temporal accuracy and trajectory realism.]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Physically Grounded Monocular Depth via Nanophotonic Wavefront Encoding</title>
		<link>https://www.immersivecomputinglab.org/publication/physically-grounded-monocular-depth-via-nanophotonic-wavefront-encoding/</link>
		
		<dc:creator><![CDATA[Bingxuan Li]]></dc:creator>
		<pubDate>Thu, 02 Jul 2026 02:49:35 +0000</pubDate>
				<category><![CDATA[Publications]]></category>
		<guid isPermaLink="false">https://www.immersivecomputinglab.org/?post_type=publication&#038;p=3803</guid>

					<description><![CDATA[Depth foundation models (DFMs) offer strong learned priors for 3D perception from single RGB images but lack physical depth cues, leading to ambiguities in metric scale. We introduce metalenses, an emerging class of ultrathin planar optical elements, as a solution to physically encode missing metric depth cues via nanophotonics. In this paper, we bridge the gap between metalens and DFMs to achieve accurate metric monocular depth sensing. In a single monocular shot, our metalens embeds depth-dependent positional shifts into two polarized optical wavefronts. With an input adaptation strategty, we enable direct fine-tuning that aligns a pretrained DFM with the optical signals. To scale the training data, we further develop a comprehensive simulation pipeline that synthesizes metalens responses from RGB-D datasets, incorporating physical factors to minimize the sim-to-real gap. Experiments demonstrate that this approach outperforms both monocular metric depth estimation and depth-from-defocus baselines, showing an effective pathway for accurate monocular metric depth sensing.]]></description>
										<content:encoded><![CDATA[Depth foundation models (DFMs) offer strong learned priors for 3D perception from single RGB images but lack physical depth cues, leading to ambiguities in metric scale. We introduce metalenses, an emerging class of ultrathin planar optical elements, as a solution to physically encode missing metric depth cues via nanophotonics. In this paper, we bridge the gap between metalens and DFMs to achieve accurate metric monocular depth sensing. In a single monocular shot, our metalens embeds depth-dependent positional shifts into two polarized optical wavefronts. With an input adaptation strategty, we enable direct fine-tuning that aligns a pretrained DFM with the optical signals. To scale the training data, we further develop a comprehensive simulation pipeline that synthesizes metalens responses from RGB-D datasets, incorporating physical factors to minimize the sim-to-real gap. Experiments demonstrate that this approach outperforms both monocular metric depth estimation and depth-from-defocus baselines, showing an effective pathway for accurate monocular metric depth sensing.]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Dichoptic Foveation</title>
		<link>https://www.immersivecomputinglab.org/publication/dichoptic-foveation/</link>
		
		<dc:creator><![CDATA[Henry Kam]]></dc:creator>
		<pubDate>Sun, 03 May 2026 20:13:00 +0000</pubDate>
				<category><![CDATA[Publications]]></category>
		<guid isPermaLink="false">https://www.immersivecomputinglab.org/?post_type=publication&#038;p=3720</guid>

					<description><![CDATA[Interocular differences in visual perception can induce a variety of effects when fused by the brain. For example, prior works have found that carefully crafted binocular differences in local detail can improve contrast. It has also been found that when the frequency content of two stimuli are slightly different, blur suppression leads to a fused percept that is typically dominated by the sharper image. In this paper, we develop a psychophysical framework to measure the perception of natural image stimuli with interocular frequency differences across the visual field. To this end, we study the effect of dichoptic foveation, which we define as the application of blur to one eye and a simultaneous sharpening filter to the other. Stimuli were viewed in a virtual reality (VR) head-mounted display (HMD) and placed at different retinal eccentricities. Study data were scaled to a perceptual just objectionable difference (JOD) scale, and a 4D model was fit to it; our results suggest that interocular frequency differences can be well described by a simple computational model. We applied the model in a realistic VR scenario with free exploration of 360° videos to improve a base foveated rendering system by enhancing high frequency information dichoptically.]]></description>
										<content:encoded><![CDATA[Interocular differences in visual perception can induce a variety of effects when fused by the brain. For example, prior works have found that carefully crafted binocular differences in local detail can improve contrast. It has also been found that when the frequency content of two stimuli are slightly different, blur suppression leads to a fused percept that is typically dominated by the sharper image. In this paper, we develop a psychophysical framework to measure the perception of natural image stimuli with interocular frequency differences across the visual field. To this end, we study the effect of dichoptic foveation, which we define as the application of blur to one eye and a simultaneous sharpening filter to the other. Stimuli were viewed in a virtual reality (VR) head-mounted display (HMD) and placed at different retinal eccentricities. Study data were scaled to a perceptual just objectionable difference (JOD) scale, and a 4D model was fit to it; our results suggest that interocular frequency differences can be well described by a simple computational model. We applied the model in a realistic VR scenario with free exploration of 360° videos to improve a base foveated rendering system by enhancing high frequency information dichoptically.]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Adapting Quality Metrics to Tone Mapping</title>
		<link>https://www.immersivecomputinglab.org/publication/adapting-quality-metrics-to-tone-mapping/</link>
		
		<dc:creator><![CDATA[Kenny Chen]]></dc:creator>
		<pubDate>Sat, 02 May 2026 20:31:04 +0000</pubDate>
				<category><![CDATA[Publications]]></category>
		<guid isPermaLink="false">https://www.immersivecomputinglab.org/?post_type=publication&#038;p=3707</guid>

					<description><![CDATA[Tone mapping evaluation is difficult because of the substantial differences in absolute luminance between high dynamic range (HDR) reference and tone-mapped standard dynamic range (SDR) test content. To address this challenge, we collected a new tone mapping evaluation dataset, focused on fundamental tone mapping operations, and combined it with several existing tone mapping quality assessment datasets. Rather than introducing new specialized metrics designed for tone-mapped content, we instead developed a set of techniques to adapt existing quality metrics for tone mapping quality assessment. Our approach models the photometric differences between HDR reference and SDR test displays for accurate metric predictions. The technique consists of two steps: first, a display model converts display-encoded content to photometric values; second, these values are re-encoded using a perceptual transfer function to map both HDR and tone-mapped images to the same display-encoded color space. We systematically evaluated both general-purpose image and video quality metrics with our adaptations and those specifically designed for tone mapping. With these adjustments, general-purpose metrics perform much better for tone mapping evaluation, consistently outperforming previously established specialized techniques. Additionally, we adapted the ColorVideoVDP metric to be sensitive to absolute luminance changes, resulting in \textit{\ourmethod}, which shows greatly improved performance and accepts photometric values as input. These results highlight the robustness of our adaptation technique and provide an improved protocol to evaluate future tone mapping quality metrics. Our datasets, code, and supplementary results can be found at kenchen10.github.io/projects/tmometric/index.html.]]></description>
										<content:encoded><![CDATA[Tone mapping evaluation is difficult because of the substantial differences in absolute luminance between high dynamic range (HDR) reference and tone-mapped standard dynamic range (SDR) test content. To address this challenge, we collected a new tone mapping evaluation dataset, focused on fundamental tone mapping operations, and combined it with several existing tone mapping quality assessment datasets. Rather than introducing new specialized metrics designed for tone-mapped content, we instead developed a set of techniques to adapt existing quality metrics for tone mapping quality assessment. Our approach models the photometric differences between HDR reference and SDR test displays for accurate metric predictions. The technique consists of two steps: first, a display model converts display-encoded content to photometric values; second, these values are re-encoded using a perceptual transfer function to map both HDR and tone-mapped images to the same display-encoded color space. We systematically evaluated both general-purpose image and video quality metrics with our adaptations and those specifically designed for tone mapping. With these adjustments, general-purpose metrics perform much better for tone mapping evaluation, consistently outperforming previously established specialized techniques. Additionally, we adapted the ColorVideoVDP metric to be sensitive to absolute luminance changes, resulting in \textit{\ourmethod}, which shows greatly improved performance and accepts photometric values as input. These results highlight the robustness of our adaptation technique and provide an improved protocol to evaluate future tone mapping quality metrics. Our datasets, code, and supplementary results can be found at kenchen10.github.io/projects/tmometric/index.html.]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Dynamic Visual Dominance in Stereoscopic Foveation</title>
		<link>https://www.immersivecomputinglab.org/publication/dynamic-visual-dominance-in-stereoscopic-foveation/</link>
		
		<dc:creator><![CDATA[Colin Groth]]></dc:creator>
		<pubDate>Fri, 01 May 2026 18:09:41 +0000</pubDate>
				<category><![CDATA[Publications]]></category>
		<guid isPermaLink="false">https://www.immersivecomputinglab.org/?post_type=publication&#038;p=3744</guid>

					<description><![CDATA[In human vision, the inputs perceived by each eye do not contribute equally to the final percept. Instead, visual dominance influences how these inputs are fused, giving more weight to one eye view over the other.  While eye dominance has been traditionally treated as a static, eye-fixed property, recent evidence suggests that dominance can vary with viewing conditions. In this work, we systematically characterize dynamic visual dominance across the visual field, with a particular focus on peripheral vision, where perceptual asymmetries are most relevant for stereoscopic rendering. Through two complementary psychophysical experiments, we first show that tolerance to eye-asymmetric blur at the fovea under binocular viewing depends on gaze direction, confirming that eye dominance is not spatially invariant. We then show that peripheral dominance is primarily governed by retinal eccentricity, with consistent naso-temporal asymmetries and dominance reversals around the blind spots. We leverage our insights in a dominance-contingent rendering application, where additional blur is selectively applied to the perceptually non-dominant eye regions under binocular viewing. Compared to static dominance approaches, our method enables stronger localized quality reductions, illustrating the practical relevance of dynamic peripheral dominance for stereoscopic foveated rendering. Thus, our goal through this work is to show how visual dominance behaves dynamically in both fovea and periphery, indicating how foveation techniques could benefit from it.]]></description>
										<content:encoded><![CDATA[In human vision, the inputs perceived by each eye do not contribute equally to the final percept. Instead, visual dominance influences how these inputs are fused, giving more weight to one eye view over the other.  While eye dominance has been traditionally treated as a static, eye-fixed property, recent evidence suggests that dominance can vary with viewing conditions. In this work, we systematically characterize dynamic visual dominance across the visual field, with a particular focus on peripheral vision, where perceptual asymmetries are most relevant for stereoscopic rendering. Through two complementary psychophysical experiments, we first show that tolerance to eye-asymmetric blur at the fovea under binocular viewing depends on gaze direction, confirming that eye dominance is not spatially invariant. We then show that peripheral dominance is primarily governed by retinal eccentricity, with consistent naso-temporal asymmetries and dominance reversals around the blind spots. We leverage our insights in a dominance-contingent rendering application, where additional blur is selectively applied to the perceptually non-dominant eye regions under binocular viewing. Compared to static dominance approaches, our method enables stronger localized quality reductions, illustrating the practical relevance of dynamic peripheral dominance for stereoscopic foveated rendering. Thus, our goal through this work is to show how visual dominance behaves dynamically in both fovea and periphery, indicating how foveation techniques could benefit from it.]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Cost-Aware Routing for Efficient Text-To-Image Generation</title>
		<link>https://www.immersivecomputinglab.org/publication/cost-aware-routing-for-efficient-text-to-image-generation/</link>
		
		<dc:creator><![CDATA[Qi Sun]]></dc:creator>
		<pubDate>Tue, 10 Mar 2026 20:46:49 +0000</pubDate>
				<category><![CDATA[Publications]]></category>
		<guid isPermaLink="false">https://www.immersivecomputinglab.org/?post_type=publication&#038;p=3676</guid>

					<description><![CDATA[Diffusion models are well known for their ability to generate a high-fidelity image for an input prompt through an iterative denoising process. Unfortunately, the high fidelity also comes at a high computational cost due to the inherently sequential generative process. In this work, we seek to optimally balance quality and computational cost, and propose a framework to allow the amount of computation to vary for each prompt, depending on its complexity. Each prompt is automatically routed to the most appropriate text-to-image generation function, which may correspond to a distinct number of denoising steps of a diffusion model, or a
disparate, independent text-to-image model. Unlike uniform cost reduction techniques (e.g., distillation, model quantization), our approach achieves the optimal trade-off by learning to reserve expensive choices (e.g., 100+ denoising steps) only for a few complex prompts, and employ more economical choices (e.g., small distilled model) for less sophisticated prompts. We empirically demonstrate on COCO and DiffusionDB that by learning to route to nine already-trained text-to-image models, our approach is able to deliver an average quality that is higher than that achievable by any of these models alone. ]]></description>
										<content:encoded><![CDATA[Diffusion models are well known for their ability to generate a high-fidelity image for an input prompt through an iterative denoising process. Unfortunately, the high fidelity also comes at a high computational cost due to the inherently sequential generative process. In this work, we seek to optimally balance quality and computational cost, and propose a framework to allow the amount of computation to vary for each prompt, depending on its complexity. Each prompt is automatically routed to the most appropriate text-to-image generation function, which may correspond to a distinct number of denoising steps of a diffusion model, or a
disparate, independent text-to-image model. Unlike uniform cost reduction techniques (e.g., distillation, model quantization), our approach achieves the optimal trade-off by learning to reserve expensive choices (e.g., 100+ denoising steps) only for a few complex prompts, and employ more economical choices (e.g., small distilled model) for less sophisticated prompts. We empirically demonstrate on COCO and DiffusionDB that by learning to route to nine already-trained text-to-image models, our approach is able to deliver an average quality that is higher than that achievable by any of these models alone. ]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>ML-PEA: Machine Learning-Based Perceptual Algorithms for Display Power Optimization</title>
		<link>https://www.immersivecomputinglab.org/publication/ml-pea-machine-learning-based-perceptual-algorithms-for-display-power-optimization/</link>
		
		<dc:creator><![CDATA[Kenny Chen]]></dc:creator>
		<pubDate>Fri, 06 Feb 2026 20:49:41 +0000</pubDate>
				<category><![CDATA[Publications]]></category>
		<guid isPermaLink="false">https://www.immersivecomputinglab.org/?post_type=publication&#038;p=3632</guid>

					<description><![CDATA[]]></description>
										<content:encoded><![CDATA[]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>HOICraft: In-Situ VLM-based Authoring Tool for Part-Level Hand-Object Interaction Design in VR</title>
		<link>https://www.immersivecomputinglab.org/publication/hoicraft-in-situ-vlm-based-authoring-tool-for-part-level-hand-object-interaction-design-in-vr/</link>
		
		<dc:creator><![CDATA[Qi Sun]]></dc:creator>
		<pubDate>Fri, 06 Feb 2026 19:54:49 +0000</pubDate>
				<category><![CDATA[Publications]]></category>
		<guid isPermaLink="false">https://www.immersivecomputinglab.org/?post_type=publication&#038;p=3625</guid>

					<description><![CDATA[]]></description>
										<content:encoded><![CDATA[]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Perceptual Impact of Peak Luminance and Contrast in Direct View HDR Display</title>
		<link>https://www.immersivecomputinglab.org/publication/perceptual-impact-of-peak-luminance-and-contrast-in-direct-view-hdr-display/</link>
		
		<dc:creator><![CDATA[Kenny Chen]]></dc:creator>
		<pubDate>Thu, 29 Jan 2026 14:35:51 +0000</pubDate>
				<category><![CDATA[Publications]]></category>
		<guid isPermaLink="false">https://www.immersivecomputinglab.org/?post_type=publication&#038;p=3619</guid>

					<description><![CDATA[]]></description>
										<content:encoded><![CDATA[]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Overdriving Visual Depth Perception via Sound Modulation in VR</title>
		<link>https://www.immersivecomputinglab.org/publication/overdriving-visual-depth-perception-via-sound-modulation-in-vr/</link>
		
		<dc:creator><![CDATA[Qi Sun]]></dc:creator>
		<pubDate>Mon, 26 Jan 2026 22:22:29 +0000</pubDate>
				<category><![CDATA[Publications]]></category>
		<guid isPermaLink="false">https://www.immersivecomputinglab.org/?post_type=publication&#038;p=3615</guid>

					<description><![CDATA[]]></description>
										<content:encoded><![CDATA[]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
