Diffusion-Aided Joint Source Channel Coding For High Realism Wireless Image Transmission
April 27, 2024 Β· Declared Dead Β· π IEEE Transactions on Machine Learning in Communications and Networking
"No code URL or promise found in abstract"
Evidence collected by the PWNC Scanner
Authors
Mingyu Yang, Bowen Liu, Boyang Wang, Hun-Seok Kim
arXiv ID
2404.17736
Category
eess.SP: Signal Processing
Cross-listed
cs.CV,
cs.IT,
eess.IV
Citations
19
Venue
IEEE Transactions on Machine Learning in Communications and Networking
Last Checked
4 months ago
Abstract
Deep learning-based joint source-channel coding (deep JSCC) has been demonstrated to be an effective approach for wireless image transmission. Nevertheless, most existing work adopts an autoencoder framework to optimize conventional criteria such as Mean Squared Error (MSE) and Structural Similarity Index (SSIM) which do not suffice to maintain the perceptual quality of reconstructed images. Such an issue is more prominent under stringent bandwidth constraints or low signal-to-noise ratio (SNR) conditions. To tackle this challenge, we propose DiffJSCC, a novel framework that leverages the prior knowledge of the pre-trained Statble Diffusion model to produce high-realism images via the conditional diffusion denoising process. Our DiffJSCC first extracts multimodal spatial and textual features from the noisy channel symbols in the generation phase. Then, it produces an initial reconstructed image as an intermediate representation to aid robust feature extraction and a stable training process. In the following diffusion step, DiffJSCC uses the derived multimodal features, together with channel state information such as the signal-to-noise ratio (SNR), as conditions to guide the denoising diffusion process, which converts the initial random noise to the final reconstruction. DiffJSCC employs a novel control module to fine-tune the Stable Diffusion model and adjust it to the multimodal conditions. Extensive experiments on diverse datasets reveal that our method significantly surpasses prior deep JSCC approaches on both perceptual metrics and downstream task performance, showcasing its ability to preserve the semantics of the original transmitted images. Notably, DiffJSCC can achieve highly realistic reconstructions for 768x512 pixel Kodak images with only 3072 symbols (<0.008 symbols per pixel) under 1dB SNR channels.
Community Contributions
Found the code? Know the venue? Think something is wrong? Let us know!
π Similar Papers
In the same crypt β Signal Processing
R.I.P.
π»
Ghosted
π
π
The Cartographer
1D Convolutional Neural Networks and Applications: A Survey
R.I.P.
π»
Ghosted
Wireless Communications with Reconfigurable Intelligent Surface: Path Loss Modeling and Experimental Measurement
π
π
The Cartographer
Accessing From The Sky: A Tutorial on UAV Communications for 5G and Beyond
R.I.P.
π»
Ghosted
6G Wireless Systems: Vision, Requirements, Challenges, Insights, and Opportunities
R.I.P.
π»
Ghosted
A New Wireless Communication Paradigm through Software-controlled Metasurfaces
Died the same way β π» Ghosted
R.I.P.
π»
Ghosted
Federated Learning: Strategies for Improving Communication Efficiency
R.I.P.
π»
Ghosted
In-Datacenter Performance Analysis of a Tensor Processing Unit
R.I.P.
π»
Ghosted
Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer Learning
R.I.P.
π»
Ghosted