跳到主要內容

[Python] 正規表達式 Regex:group

記錄正規表達式的 group 擷取字串的順序。
Python 的正規表達式,透過 group() 函式擷取括號框住的內容。

以下是我寫的例子:
import re
def main():
    string_to_search_1 = "food   @123"
    string_to_search_2 = "feed"
    pattern = re.compile(r'(f(oo|ee)d)\s*(@(\d+))*')

    matches = pattern.finditer(string_to_search_1)
    for match in matches:
        print "match:", match.group()
        print "1:", match.group(1)
        print "2:", match.group(2)
        print "3:", match.group(3)
        print "4:", match.group(4)
        # print "5:", match.group(5)  # no such group.

        tmp = int(match.group(4))
        tmp += 5
        print "tmp =", tmp
        tmp = str(match.group(4))
        print "tmp[2] =", tmp[2]

    matches = pattern.finditer(string_to_search_2)
    for match in matches:
        print "match:", match.group()
        print "1:", match.group(1)
        print "2:", match.group(2)
        print "3:", match.group(3)  #
        print "4:", match.group(4)  #
        # print "5:", match.group(5)  # no such group.

        if match.group(4) is None: print " None detected in group(4)."
        tmp = str(match.group(4))
        print "tmp:", tmp

if(__name__ == "__main__"):
    main()


執行結果
match: food   @123
1: food
2: oo
3: @123
4: 123
tmp = 128
tmp[2] = 3
match: feed
1: feed
2: ee
3: None
4: None
 None detected in group(4).
tmp: None

幾個重點
  1. 在正規表達式內,有幾個括號,就有幾組 group。
  2. 在 string_to_search_1 內掃描 (f(oo|ee)d) 這個 Pattern,group(1) 會回傳 food,group(2) 回傳 "oo"。先從最外大括號先回傳,其次再回傳內部小括號的字串。
  3. 回傳的字串可以自由的轉成整數或是取出個別字元。
  4. 在 string_to_search_2 內,group(3) 跟 group(4) 都回傳 None,因為找不到對應 Pattern。

留言

這個網誌中的熱門文章

[Linux] Elementary OS 字體調校

用 gesetting 取得 elementary os 的等寬字體(這也是終端機默認字體): gsettings get org.gnome.desktop.interface monospace-font-name 會顯示目前字型跟字體大小: Roboto Mono 10 設定字體大小: gsettings set org.gnome.desktop.interface monospace-font-name 'Roboto Mono 12' 可以微調 text-scaling-factor: gsettings set org.gnome.desktop.interface text-scaling-factor <value>

[心得] 復古、老派的程式設計之路

最近看到 John Carmack 2018 年的貼文中的幾段話: I’m not a Unix geek.  I get around ok, but I am most comfortable developing in Visual Studio on Windows.  I thought a week of full immersion work in the old school Unix style would be interesting, even if it meant working at a slower pace.  It was sort of an adventure in retro computing — this was fvwm and vi.  Not vim, actual BSD vi. In the end, I didn’t really explore the system all that much, with 95% of my time in just the basic vi / make / gdb operations.  I appreciated the good man pages, as I tried to do everything within the self contained system, without resorting to internet searches.  Seeing references to 30+ year old things like Tektronix terminals was amusing. In the spirit of my retro theme, I had printed out several of Yann LeCun’s old papers and was considering doing everything completely off line, as if I was actually in a mountain cabin somewhere, but I wound up watching a lot of the Stanford CS231N lectur...

[C++] 用 stringstream 分割同時包含 ASCII 和 Unicode 中文的字串

假設有一個檔案,內容包含中文,也包含 ASCII 字元(皆是可顯示字元)。 如果這個檔案格式為 UTF-8,則 ASCII 碼佔用 1 個 byte,中文字佔用 3 個 byte。且代表中文的這 3 個 byte,必定是負值。 (其他語言的文字,未必是用 3 byte 儲存,但繁體中文佔用了 3 byte,UTF-8 的特性就是 byte 數是變動的) 所以 stringstream 用 UTF-8 的 ASCII 的可顯示字元,來當作 delimiter分割字串,是很安全的,例如用 ASCII 的 "空白" 來將以下的檔案,每一行輸入都拆解成兩個獨立字串: apple 蘋果 橘子 哈哈 兔子 rabbit 可以這樣寫: ifstream inputFile(fileName.c_str());  string a, b, line; while(getline(inputFile, line)) {     stringstream tkn(line);     tkn >> a >> b;         cout << " [Debug] a == " << a << endl;     cout << " [Debug] b == " << b << endl; } inputFile.close(); 輸出會是: [Debug] a == apple [Debug] b == 蘋果 [Debug] a == 橘子 [Debug] b == 哈哈 [Debug] a == 兔子 [Debug] b == rabbit 一般我很少直接採用 Unicode 編碼來處理字串,較常採用 UTF-8 。因為 UTF-8 的前 128 個編碼,跟 ASCII 編碼是完全相同的,如此方便程式碼可以相容 ASCII 編碼,同時又可以處理寬字元。許多支持多國語言的程式,預設都用 UTF-8 就是這個緣故。 注意在 Window 系統記事本若將檔案存成 UT...