IMAGE AND VISION COMPUTING图像与视觉计算

IMAGE AND VISION COMPUTING(英文缩写 IMAGE VISION COMPUT),ISSN 0262-8856,eISSN 1872-8138,中文译名:图像与视觉计算 是一本学术期刊。本页汇总该期刊的最新影响因子、分区信息以及最新收录于 PubMed 的文献,帮助您快速了解期刊全貌。

2026 年数据 · 影响因子
5.000
JCR 分区
Q1
CAS 分区
B3
近一年发文量
0
本站 PubMed 收录统计

发文量统计区间:2025-09-27 至 2026-09-27,按本站收录文献的发表日期统计。

ISSN: 0262-8856 · eISSN: 1872-8138 · 缩写: IMAGE VISION COMPUT ·中文: 图像与视觉计算

期刊介绍

选择期刊介绍栏目

期刊简介

《Image and Vision Computing》是一本国际同行评议期刊,聚焦计算机视觉与图像理解的理论与应用研究。主要领域涵盖图像处理、模式识别、三维视觉、视频分析、机器学习在视觉中的应用等。读者群包括计算机科学、工程和认知科学领域的研究人员、工程师及研究生,尤其适合关注视觉算法创新与实际系统实现的群体。

研究方向

期刊主要发表图像与视觉计算方向的高质量原创论文,主题包括图像分割、目标检测与跟踪、人脸与生物识别、场景理解、立体视觉、运动分析、深度学习模型及其视觉应用。论文类型以理论方法、算法创新和实验验证为主,也接受综述和简短通讯。

期刊特色

研究取向强调方法新颖性与实验充分性,鼓励跨学科融合和实际场景验证。论文通常包含算法推导、对比实验和公开数据集评估。适合计算机视觉、图像处理及相关领域的研究者、博士研究生和工业界研发人员阅读与投稿。

投稿难度

投稿难度中等偏上,期刊对方法创新和实验严谨性要求较高。建议在投稿前确保算法对比全面、数据集多样,并清晰阐述与已有工作的差异。即使分区指标不突出,也不应低估其同行评议的严格性,需认真准备回复审稿意见。

历年影响因子趋势

JCR 数据年份影响因子JCR 分区
20213.860Q1
20224.700Q1
20234.200Q1
20244.200Q1
20255.000Q1

IMAGE AND VISION COMPUTING 最新收录文献

  1. JCR分区: Q1 CAS分区: B3 影响因子: 5

    1. A survey on computer vision based human analysis in the COVID-19 era.

    1. COVID-19时代基于计算机视觉的人体分析综述
    作者:
    Fevziye Irem Eyiokur, Alperen Kantarcı, Mustafa Ekrem Erakın, Naser Damer, Ferda Ofli, Muhammad Imran, Janez Križaj, Albert Ali Salah, Alexander Waibel, Vitomir Štruc, Hazım Kemal Ekenel
    日期:
    2023-02-01

    The emergence of COVID-19 has had a global and profound impact, not only on society as a whole, but also on the lives of individuals. Various prevention measures were introduced around the world to limit the transmission of the disease, including face masks, mandates for social distancing and regular disinfection in public spaces, and the use of screening applications. These developments also triggered the need for novel and improved computer vision techniques capable of providing support to the prevention measures through an automated analysis of visual data, on the one hand, and facilitating normal operation of existing vision-based services, such as biometric authentication schemes, on the other. Especially important here, are computer vision techniques that focus on the analysis of people and faces in visual data and have been affected the most by the partial occlusions introduced by the mandates for facial masks. Such computer vision based human analysis techniques include face and face-mask detection approaches, face recognition techniques, crowd counting solutions, age and expression estimation procedures, models for detecting face-hand interactions and many others, and have seen considerable attention over recent years. The goal of this survey is to provide an introduction to the problems induced by COVID-19 into such research and to present a comprehensive review of the work done in the computer vision based human analysis field. Particular attention is paid to the impact of facial masks on the performance of various methods and recent solutions to mitigate this problem. Additionally, a detailed review of existing datasets useful for the development and evaluation of methods for COVID-19 related applications is also provided. Finally, to help advance the field further, a discussion on the main open challenges and future research direction is given at the end of the survey. This work is intended to have a broad appeal and be useful not only for computer vision researchers but also the general public.

  2. JCR分区: Q1 CAS分区: B3 影响因子: 5

    2. Progressive ShallowNet for large scale dynamic and spontaneous facial behaviour analysis in children.

    作者:
    Abdul Qayyum, Imran Razzak, Nour Moustafa, Moona Mazher
    日期:
    2022-03-01

    COVID-19 has severely disrupted every aspect of society and left negative impact on our life. Resisting the temptation in engaging face-to-face social connection is not as easy as we imagine. Breaking ties within social circle makes us lonely and isolated, that in turns increase the likelihood of depression related disease and even can leads to death by increasing the chance of heart disease. Not only adults, children's are equally impacted where the contribution of emotional competence to social competence has long term implications. Early identification skill for facial behaviour emotions, deficits, and expression may help to prevent the low social functioning. Deficits in young children's ability to differentiate human emotions can leads to social functioning impairment. However, the existing work focus on adult emotions recognition mostly and ignores emotion recognition in children. By considering the working of pyramidal cells in the cerebral cortex, in this paper, we present progressive lightweight shallow learning for the classification by efficiently utilizing the skip-connection for spontaneous facial behaviour recognition in children. Unlike earlier deep neural networks, we limit the alternative path for the gradient at the earlier part of the network by increase gradually with the depth of the network. Progressive ShallowNet is not only able to explore more feature space but also resolve the over-fitting issue for smaller data, due to limiting the residual path locally, making the network vulnerable to perturbations. We have conducted extensive experiments on benchmark facial behaviour analysis in children that showed significant performance gain comparatively.

  3. JCR分区: Q1 CAS分区: B3 影响因子: 5

    3. FMD-Yolo: An efficient face mask detection method for COVID-19 prevention and control in public.

    作者:
    Peishu Wu, Han Li, Nianyin Zeng, Fengping Li
    日期:
    2022-01-01

    Coronavirus disease 2019 (COVID-19) is a world-wide epidemic and efficient prevention and control of this disease has become the focus of global scientific communities. In this paper, a novel face mask detection framework FMD-Yolo is proposed to monitor whether people wear masks in a right way in public, which is an effective way to block the virus transmission. In particular, the feature extractor employs Im-Res2Net-101 which combines Res2Net module and deep residual network, where utilization of hierarchical convolutional structure, deformable convolution and non-local mechanisms enables thorough information extraction from the input. Afterwards, an enhanced path aggregation network En-PAN is applied for feature fusion, where high-level semantic information and low-level details are sufficiently merged so that the model robustness and generalization ability can be enhanced. Moreover, localization loss is designed and adopted in model training phase, and Matrix NMS method is used in the inference stage to improve the detection efficiency and accuracy. Benchmark evaluation is performed on two public databases with the results compared with other eight state-of-the-art detection algorithms. At = 0.5 level, proposed FMD-Yolo has achieved the best precision 50 of 92.0% and 88.4% on the two datasets, and 75 at = 0.75 has improved 5.5% and 3.9% respectively compared with the second one, which demonstrates the superiority of FMD-Yolo in face mask detection with both theoretical values and practical significance.

  4. JCR分区: Q1 CAS分区: B3 影响因子: 5

    4. Learning Facial Action Units with Spatiotemporal Cues and Multi-label Sampling.

    作者:
    Wen-Sheng Chu, Fernando De la Torre, Jeffrey F Cohn
    日期:
    2019-01-01

    Facial action units (AUs) may be represented , and in terms of their . Previous research focuses on one or another of these aspects or addresses them disjointly. We propose a hybrid network architecture that jointly models spatial and temporal representations and their correlation. In particular, we use a Convolutional Neural Network (CNN) to learn spatial representations, and a Long Short-Term Memory (LSTM) to model temporal dependencies among them. The outputs of CNNs and LSTMs are aggregated into a fusion network to produce per-frame prediction of multiple AUs. The hybrid network was compared to previous state-of-the-art approaches in two large FACS-coded video databases, GFT and BP4D, with over 400,000 AU-coded frames of spontaneous facial behavior in varied social contexts. Relative to standard multi-label CNN and feature-based state-of-the-art approaches, the hybrid system reduced person-specific biases and obtained increased accuracy for AU detection. To address class imbalance within and between batches during training the network, we introduce multi-labeling sampling strategies that further increase accuracy when AUs are relatively sparse. Finally, we provide visualization of the learned AU models, which, to the best of our best knowledge, reveal for the first time how machines see AUs.

  5. JCR分区: Q1 CAS分区: B3 影响因子: 5

    5. Lessons from Collecting a Million Biometric Samples.

    作者:
    P Jonathon Phillips, Patrick J Flynn, Kevin W Bowyer
    日期:
    2017-02-01

    Over the past decade, independent evaluations have become commonplace in many areas of experimental computer science, including face and gesture recognition. A key attribute of many successful independent evaluations is a curated data set. Desired aspects associated with these data sets include appropriateness to the experimental design, a corpus size large enough to allow statistically rigorous characterization of results, and the availability of comprehensive metadata that allow inferences to be made on various data set attributes. In this paper, we review a ten-year biometric sampling effort that enabled the creation of several key biometrics challenge problems. We summarize the design and execution of data collections, identify key challenges, and convey some lessons learned.

  6. JCR分区: Q1 CAS分区: B3 影响因子: 5

    6. Dense 3D Face Alignment from 2D Video for Real-Time Use.

    作者:
    László A Jeni, Jeffrey F Cohn, Takeo Kanade
    日期:
    2017-02-01

    To enable real-time, person-independent 3D registration from 2D video, we developed a 3D cascade regression approach in which facial landmarks remain invariant across pose over a range of approximately 60 degrees. From a single 2D image of a person's face, a dense 3D shape is registered in real time for each frame. The algorithm utilizes a fast cascade regression framework trained on high-resolution 3D face-scans of posed and spontaneous emotion expression. The algorithm first estimates the location of a dense set of landmarks and their visibility, then reconstructs face shapes by fitting a part-based 3D model. Because no assumptions are required about illumination or surface properties, the method can be applied to a wide range of imaging conditions that include 2D video and uncalibrated multi-view video. The method has been validated in a battery of experiments that evaluate its precision of 3D reconstruction, extension to multi-view reconstruction, temporal integration for videos and 3D head-pose estimation. Experimental findings strongly support the validity of real-time, 3D registration and reconstruction from 2D video. The software is available online at http://zface.org.

  7. JCR分区: Q1 CAS分区: B3 影响因子: 5

    7. Multiview stereo and silhouette fusion via minimizing generalized reprojection error.

    作者:
    Zhaoxin Li, Kuanquan Wang, Wenyan Jia, Hsin-Chen Chen, Wangmeng Zuo, Deyu Meng, Mingui Sun
    日期:
    2015-01-01

    Accurate reconstruction of 3D geometrical shape from a set of calibrated 2D multiview images is an active yet challenging task in computer vision. The existing multiview stereo methods usually perform poorly in recovering deeply concave and thinly protruding structures, and suffer from several common problems like slow convergence, sensitivity to initial conditions, and high memory requirements. To address these issues, we propose a two-phase optimization method for generalized reprojection error minimization (TwGREM), where a generalized framework of reprojection error is proposed to integrate stereo and silhouette cues into a unified energy function. For the minimization of the function, we first introduce a convex relaxation on 3D volumetric grids which can be efficiently solved using variable splitting and Chambolle projection. Then, the resulting surface is parameterized as a triangle mesh and refined using surface evolution to obtain a high-quality 3D reconstruction. Our comparative experiments with several state-of-the-art methods show that the performance of TwGREM based 3D reconstruction is among the highest with respect to accuracy and efficiency, especially for data with smooth texture and sparsely sampled viewpoints.

  8. JCR分区: Q1 CAS分区: B3 影响因子: 5

    8. Nonverbal Social Withdrawal in Depression: Evidence from manual and automatic analysis.

    作者:
    Jeffrey M Girard, Jeffrey F Cohn, Mohammad H Mahoor, S Mohammad Mavadati, Zakia Hammal, Dean P Rosenwald
    日期:
    2014-10-01

    The relationship between nonverbal behavior and severity of depression was investigated by following depressed participants over the course of treatment and video recording a series of clinical interviews. Facial expressions and head pose were analyzed from video using manual and automatic systems. Both systems were highly consistent for FACS action units (AUs) and showed similar effects for change over time in depression severity. When symptom severity was high, participants made fewer affiliative facial expressions (AUs 12 and 15) and more non-affiliative facial expressions (AU 14). Participants also exhibited diminished head motion (i.e., amplitude and velocity) when symptom severity was high. These results are consistent with the Social Withdrawal hypothesis: that depressed individuals use nonverbal behavior to maintain or increase interpersonal distance. As individuals recover, they send more signals indicating a willingness to affiliate. The finding that automatic facial expression analysis was both consistent with manual coding and revealed the same pattern of findings suggests that automatic facial expression analysis may be ready to relieve the burden of manual coding in behavioral and clinical science.

  9. JCR分区: Q1 CAS分区: B3 影响因子: 5

    9. Classification and Weakly Supervised Pain Localization using Multiple Segment Representation.

    作者:
    Karan Sikka, Abhinav Dhall, Marian Stewart Bartlett
    日期:
    2014-10-01

    Automatic pain recognition from videos is a vital clinical application and, owing to its spontaneous nature, poses interesting challenges to automatic facial expression recognition (AFER) research. Previous pain vs no-pain systems have highlighted two major challenges: (1) ground truth is provided for the sequence, but the presence or absence of the target expression for a given frame is unknown, and (2) the time point and the duration of the pain expression event(s) in each video are unknown. To address these issues we propose a novel framework (referred to as MS-MIL) where each sequence is represented as a bag containing multiple segments, and multiple instance learning (MIL) is employed to handle this weakly labeled data in the form of sequence level ground-truth. These segments are generated via multiple clustering of a sequence or running a multi-scale temporal scanning window, and are represented using a state-of-the-art Bag of Words (BoW) representation. This work extends the idea of detecting facial expressions through 'concept frames' to 'concept segments' and argues through extensive experiments that algorithms such as MIL are needed to reap the benefits of such representation. The key advantages of our approach are: (1) joint detection and localization of painful frames using only sequence-level ground-truth, (2) incorporation of temporal dynamics by representing the data not as individual frames but as segments, and (3) extraction of multiple segments, which is well suited to signals with uncertain temporal location and duration in the video. Extensive experiments on UNBC-McMaster Shoulder Pain dataset highlight the effectiveness of the approach by achieving competitive results on both tasks of pain classification and localization in videos. We also empirically evaluate the contributions of different components of MS-MIL. The paper also includes the visualization of discriminative facial patches, important for pain detection, as discovered by our algorithm and relates them to Action Units that have been associated with pain expression. We conclude the paper by demonstrating that MS-MIL yields a significant improvement on another spontaneous facial expression dataset, the FEEDTUM dataset.

  10. JCR分区: Q1 CAS分区: B3 影响因子: 5

    10. Non-rigid Face Tracking with Local Appearance Consistency Constraint.

    作者:
    Yang Wang, Simon Lucey, Jeffrey F Cohn, Jason Saragih
    日期:
    2010-05-01

    In this paper we present a new discriminative approach to achieve consistent and efficient tracking of non-rigid object motion, such as facial expressions. By utilizing both spatial and temporal appearance coherence at the patch level, the proposed approach can reduce ambiguity and increase accuracy. Recent research demonstrates that feature based approaches, such as constrained local models (CLMs), can achieve good performance in non-rigid object alignment/tracking using local region descriptors and a non-rigid shape prior. However, the matching performance of the learned generic patch experts is susceptible to local appearance ambiguity. Since there is no motion continuity constraint between neighboring frames of the same sequence, the resultant object alignment might not be consistent from frame to frame and the motion field is not temporally smooth. In this paper, we extend the CLM method into the spatio-temporal domain by enforcing the appearance consistency constraint of each local patch between neighboring frames. More importantly, we show that the global warp update can be optimized jointly in an efficient manner using convex quadratic fitting. Finally, we demonstrate that our approach receives improved performance for the task of non-rigid facial motion tracking on the videos of clinical patients.

在 IMAGE AND VISION COMPUTING 中搜索更多文献

支持中英文检索 · 智能翻译 · 影响因子 · PDF 下载 · AI 文献阅读

指标接近的期刊