A Probabilistic Consistency and Profile-Based Framework for Protein Multiple Sequence Alignment

Multiple Sequence Alignment (MSA) is a crucial bioinformatics task for locating conserved regions, evolutionary relationships, and functional similarities among biological sequences. The accurate alignment of proteins continues to be difficult because of factors such as sequence divergence, insertions and deletions, the positioning of gaps, and the extensive search space involved in the alignment process. A probabilistic consistency and profile-based framework for protein MSA is introduced in this study. It integrates pair-HMM posterior probabilities, posterior consistency, guide-tree construction, expected-accuracy profile alignment, and progressive profile merging. Initially, the framework builds a preliminary alignment via progressive profile alignment and assesses residue correspondences based on posterior probabilities. The proposed method was assessed on the BAliBASE RV11 and RV12 benchmarks using Sum-of-Pairs (SP) and Total-Column (TC) scores with FastSP. It reached average SP/TC scores of 0.6751/0.3644 on RV11 and 0.9096/0.6794 on RV12, with mean alignment runtimes of 1.44 and 1.87 seconds, respectively. The outcomes show a competitive accuracy for pairwise alignment while also highlighting chances for improving complete-column recovery.