Login
用户名
密码
Login
用户名
密码
周永锋课题组
Yong-feng Zhou lab
快捷键为`/~(Backquote)

JIPB近日在线发表了中国农业科学院深圳农业基因组研究所商连光课题组联合多家单位撰写的题为A centromere map based on super pan-genome highlights the structure and function of rice centromeres的研究论文https://doi.org/10.1111/jipb.13607,该研究基于水稻超级泛基因组构建了目前群体规模最大的水稻泛着丝粒图谱,为水稻着丝粒进化模式和着丝粒基因挖掘提供了新见解。

着丝粒是真核生物染色体的重要功能结构,在细胞有丝分裂和减数分裂过程中,着丝粒保障染色体正确分离和传递Cheng et al., 2002。大多数植物的着丝粒由卫星重复序列、转座子和少数基因组成,高度复杂的重复序列对其精确组装和功能解析带来了挑战,极大阻碍了对着丝粒结构、特征和功能的研究。目前,对植物着丝粒基因组学的研究主要集中于特定个体或小群体规模,缺乏对该区域的大规模群体分析,这对于全面了解它们的结构、进化和功能至关重要。因此,构建高质量的泛着丝粒图谱对水稻着丝粒区域的生物学功能研究具有重要意义,包括细胞分裂、作物驯化、重组抑制、新基因的出现等等。

Figure 1. The centromere map of 251 rice mini‐core germplasm accessions (A) Schematic representation of the construction of the rice centromere map. The pink region on the chromosome indicates the CentO‐enriched region (CoER), and the blue region on the chromosome indicates the peri‐CoER. (B–D) Tukey's multiple comparisons test of the CoER length (B) and the number of CoER elements with CentO satellite (C) and long terminal repeats (LTRs) (D) across different subpopulations in the complete CoER dataset of the centromere map. Significance is represented in alphabetical notation (P < 0.05). (E) Landscape of CoER length (i), CentO repeat copy (ii) and LTR number (iii) in the complete CoERs across different chromosomes. Ob, Og, Or, Osi, Osj, As, Af and All refer to Oryza barthii, O. glaberrima, O. rufipogon, O. sativa indica, O. sativa japonica, African rice, Asian rice, and the 251 accessions, respectively.

Figure 2. High diversity analysis of CentO repeats in the rice centromere map (A) CentO repeat copy across all complete CentO‐enriched regions (CoERs), with red bar highlighting the 155 and 165 bp CentO satellite repeats. (B) Heatmap showing the identity of CentO consensus sequences on each CoER between centromeres. With light red (lowest) to dark red (highest), based on the average similarity ratio (%) of the consensus sequences on each chromosome for each complete CoER. (C) Maximum likelihood phylogenetic tree of CentO consensus repeats from 2,188 complete CoERs. Tracks from outer to inner circles indicate: subpopulation distribution, centromere distribution of different chromosomes. (D) Comparison of CentO satellite repeats from two clades. (i) Two bar charts showing the distribution of different lengths of CentO repeats in each subpopulation within Clade I and Clade II. The x‐axis shows the CentO monomer length, and the y‐axis shows the ratio of CentO repeats with that subpopulation. (ii) Two pie charts showing the ratio of different lengths of CentO repeats within Clade I and Clade II. The CentO repeat lengths category is shown in the legend. Excluded were the CentO repeat lengths that are <10% of the total. (E) Frequency of CentO sequence variant sites across the consensus repeats from the CoERs. The color‐coded DNA sequence is shown on left (dark green, A; red, T; orange, C; light green, G; ‐, gap).

Figure 3. Characterization of long terminal repeats (LTRs) in the rice centromere map (A, B) The linear regression curves show large correlation between the number of LTRs (y‐axis) and the distance from the centromere (x‐axis) (except for some accessions on chromosomes 9 and 11, P‐values are < 2.2e‐16 by Spearman correlation analysis) (A), and the number of LTRs within CentO‐enriched regions (CoERs) (y‐axis) and the CoER length (x‐axis) (B). The linear regression curves for each subpopulation are shown as R‐values and P‐values in the legend. (C) Number variation of LTR elements in complete CoERs dataset across different subpopulations and chromosomes. The x, y, and z axes represent the mean, SD, and coefficient of variation (CV) of LTR numbers, respectively. (D) The genome‐wide distribution of LTRs in complete CoERs across different subpopulations. The x‐axis represents the distance from the CoER origin to the flanking regions (200 kb) on either side. The y‐axis represents the LTR density, and was calculated by dividing the number of LTRs in the CoER length. The subplot that shows the LTR distribution patterns in the 5 Mb flanking regions of CoER (above the main plot). (E) The LTR distribution patterns with a representative example of the CoERs on chromosome 1. Upper panel, density plot of LTR numbers in CoER start sites and genomic regions on chromosome 1. Lower panel, heatmap of LTR number distribution in each CoER start site and genomic regions on chromosome 1. The red dashed line indicates the CoERs. The window size is 200 kb. The heatmap shows the LTR density between different genome regions on chromosome 1 for each CoER, ranging from blue (lowest) to red (highest). The color‐coded representation indicates different subpopulations of accessions. (F, G) The percentage of LTRs with different insertion times (F) and different LTR families (G) in different genomic regions across complete CoERs. The regions are defined as the core centromere region (CoER) (i), the 1 Mb flanking region of the CoER on either side (peri‐CoER) (ii), the rest of the genome (non‐(peri)CoER) (iii), and genome‐wide (iv). Ob, Og, Or, Osi, and Osj refer to Oryza barthii, O. glaberrima, O. rufipogon, O. sativa indica, and O. sativa japonica, respectively.

Figure 4. Identification of a centromeric gene OsMAB based on structural variation – expression quantitative trait loci (SV‐eQTL) analysis (A) Local Manhattan plot of OsMAB expression level and SV. The gray region in the Manhattan plot indicates centromere region, and the centromere position information was obtained from the CENH3 (centromere‐specific histone H3‐like protein) chromatin immunoprecipitation ‐ sequencing of the telomere‐to‐telomere (T2T) Nipponbare genome. The black dashed line represents P‐value threshold (P‐value = 2.48e‐5). (B, C) Fragments per kilobase of transcript per million mapped reads (FPKM) (B) of OsMAB and tiller numbers (C) of accessions with (Hap.2) or without (Hap.1) and the deletion in the promoter of OsMAB. **P < 0.01, ***P < 0.001 by Wilcox test (B, C). (D) Morphology of NIP and osmab mutant. Bar, 10 cm. (E) Sequencing comparison of mutation sites in the wild type and the osmab mutant. (F) Gene structure and variation sites within coding sequences (CDs) of the OsMAB. (G) Tiller numbers of NIP compared with osmab mutant (n = 12). ***P < 0.001 by Student's t‐test. (H) Subcellular localization of OsMAB (bar, 5 μm).

该研究基于水稻完整参考基因组Shang et al., 2023 和251份水稻的超级泛基因组Shang et al., 2022构建了一个高质量的着丝粒图谱,首次在群体水平上揭示了水稻着丝粒重复序列长度和序列的多样性, 分析了着丝粒重复序列在水稻亚群之间的特征,进一步探究了水稻着丝粒在不同染色体间可能具有不同的进化模式。在群体水平上分析了水稻着丝粒区域LTR的多样性,Gypsy LTR的基因组分布模式以及其在着丝粒的形成和进化中起到了关键作用,尤其是年轻的Gypsy LTR。结合泛基因组数据,在12号染色体的着丝粒区发现了一个新基因OsMAB,该基因正向调节水稻分蘖数,并且使用CRISPR/Cas9进一步验证了这一发现。

总的来说,该研究构建大规模的、高质量的水稻泛着丝粒图谱,为理解水稻着丝粒的结构、进化和功能提供了重要依据。

中国农业科学院深圳农业基因组研究所商连光研究员、中国水稻研究所郭龙彪研究员和崖州湾国家实验室钱前院士为论文的共同通讯作者。博士研究生吕阳、硕士研究生刘聪聪、博士研究生李笑霞和王月影为论文的共同第一作者。该研究得到国家自然科学基金基础科学中心、广东省自然科学基金杰出青年基金、中国农科院青年创新专项资金资助。






图片无法显示
图片无法显示