如何从多个文件中提取特定信息并在linux中制作table?
How to extract specific information from multiple files and make a table in linux?
我有多个包含信息的文本文件。这里我展示了两个文本文件,如下所示:
Sample1.txt
Status /documents/Sample1.sorted.bam
Assigned 50945040
Unassigned_Unmapped 947866
Unassigned_MappingQuality 0
Unassigned_Chimera 0
Unassigned_FragmentLength 0
Unassigned_Duplicate 0
Unassigned_MultiMapping 49013681
Unassigned_Secondary 0
Unassigned_Nonjunction 0
Unassigned_NoFeatures 21189312
Unassigned_Overlapping_Length 0
Unassigned_Ambiguity 4430011
Sample2.txt
Status /documents/Sample2.sorted.bam
Assigned 36335614
Unassigned_Unmapped 870456
Unassigned_MappingQuality 0
Unassigned_Chimera 0
Unassigned_FragmentLength 0
Unassigned_Duplicate 0
Unassigned_MultiMapping 68688141
Unassigned_Secondary 0
Unassigned_Nonjunction 0
Unassigned_NoFeatures 23746485
Unassigned_Overlapping_Length 0
Unassigned_Ambiguity 3734593
对于单个文本文件,我使用 grep:
grep "Assigned\|Unmapped\|MultiMapping\|NoFeatures\|Ambiguity" Sample1.txt > output.txt
但我希望输出如下所示,我可以在所有文本文件上使用一个小脚本并使 table:
Sample1 Sample2
Assigned 50945040 36335614
Unassigned_Unmapped 947866 870456
Unassigned_MultiMapping 49013681 68688141
Unassigned_NoFeatures 21189312 23746485
Unassigned_Ambiguity 4430011 3734593
$ cat tst.awk
!= 0 {
printf "%s%s", (NR>1 ? : "Name"), OFS
for (i=2; i<=NF; i+=2) {
gsub(/^.*\/|\..*$/,"",$i)
printf "%s%s", $i, (i<NF ? OFS : ORS)
}
}
$ paste Sample1.txt Sample2.txt | awk -f tst.awk | column -t
Name Sample1 Sample2
Assigned 50945040 36335614
Unassigned_Unmapped 947866 870456
Unassigned_MultiMapping 49013681 68688141
Unassigned_NoFeatures 21189312 23746485
Unassigned_Ambiguity 4430011 3734593
要获得 Excel 可以理解的输出而不是问题中显示的输出,请执行以下操作:
$ cat tst.awk
BEGIN { OFS="," }
!= 0 {
printf "%s%s", (NR>1 ? : "Name"), OFS
for (i=2; i<=NF; i+=2) {
gsub(/^.*\/|\..*$/,"",$i)
printf "%s%s", $i, (i<NF ? OFS : ORS)
}
}
$ paste Sample1.txt Sample2.txt | awk -f tst.awk > output.csv
然后双击 output.csv 用 Excel 打开它。
我有多个包含信息的文本文件。这里我展示了两个文本文件,如下所示:
Sample1.txt
Status /documents/Sample1.sorted.bam
Assigned 50945040
Unassigned_Unmapped 947866
Unassigned_MappingQuality 0
Unassigned_Chimera 0
Unassigned_FragmentLength 0
Unassigned_Duplicate 0
Unassigned_MultiMapping 49013681
Unassigned_Secondary 0
Unassigned_Nonjunction 0
Unassigned_NoFeatures 21189312
Unassigned_Overlapping_Length 0
Unassigned_Ambiguity 4430011
Sample2.txt
Status /documents/Sample2.sorted.bam
Assigned 36335614
Unassigned_Unmapped 870456
Unassigned_MappingQuality 0
Unassigned_Chimera 0
Unassigned_FragmentLength 0
Unassigned_Duplicate 0
Unassigned_MultiMapping 68688141
Unassigned_Secondary 0
Unassigned_Nonjunction 0
Unassigned_NoFeatures 23746485
Unassigned_Overlapping_Length 0
Unassigned_Ambiguity 3734593
对于单个文本文件,我使用 grep:
grep "Assigned\|Unmapped\|MultiMapping\|NoFeatures\|Ambiguity" Sample1.txt > output.txt
但我希望输出如下所示,我可以在所有文本文件上使用一个小脚本并使 table:
Sample1 Sample2
Assigned 50945040 36335614
Unassigned_Unmapped 947866 870456
Unassigned_MultiMapping 49013681 68688141
Unassigned_NoFeatures 21189312 23746485
Unassigned_Ambiguity 4430011 3734593
$ cat tst.awk
!= 0 {
printf "%s%s", (NR>1 ? : "Name"), OFS
for (i=2; i<=NF; i+=2) {
gsub(/^.*\/|\..*$/,"",$i)
printf "%s%s", $i, (i<NF ? OFS : ORS)
}
}
$ paste Sample1.txt Sample2.txt | awk -f tst.awk | column -t
Name Sample1 Sample2
Assigned 50945040 36335614
Unassigned_Unmapped 947866 870456
Unassigned_MultiMapping 49013681 68688141
Unassigned_NoFeatures 21189312 23746485
Unassigned_Ambiguity 4430011 3734593
要获得 Excel 可以理解的输出而不是问题中显示的输出,请执行以下操作:
$ cat tst.awk
BEGIN { OFS="," }
!= 0 {
printf "%s%s", (NR>1 ? : "Name"), OFS
for (i=2; i<=NF; i+=2) {
gsub(/^.*\/|\..*$/,"",$i)
printf "%s%s", $i, (i<NF ? OFS : ORS)
}
}
$ paste Sample1.txt Sample2.txt | awk -f tst.awk > output.csv
然后双击 output.csv 用 Excel 打开它。