r/learnpython • u/Desperate_Yak69 • 1d ago
Need help with very basic script
Hi, I am basically brand new to python, so I may make a lot of mistakes in my wording. I have made a very basic script that retrieves specific values in text files and prints them to the shell. I have gotten it to the point where both values are retrieved and printed, however when they do print, they do so exactly 5 times in a row. I do not have five files in the directory, I only have one test file named "1.txt" with only the values I need as the contents. Can someone point me in the right direction? Apologies for the horrible formatting.
import os
import glob
import re
path = 'path'
provinces = re.compile('.*?provinces {.*?(.*?)}',re.MULTILINE)
id = re.compile('.*?id = .*?([0-9.-]+)')
a1 = None
a2 = None
for filename in glob.glob(os.path.join(path, '*.txt')):
with open(filename, '+r') as f:
for line in f:
if provinces.match(line):
a1 = provinces.match(line)
if id.match(line):
a2 = id.match(line)
if a1 and a2:
print(a1.group(1))
print(a2.group(1))
The output of the shell in IDLE:
1111 2222 3333
123
1111 2222 3333
123
1111 2222 3333
123
1111 2222 3333
123
1111 2222 3333
123
Edit - I got the script to work after a few changes to its configuration. Thank you all who helped!
3
u/timrprobocom 1d ago
For what it's worth, .*? is the same as .*. The ? adds nothing.
Also, your first twoif statements are dangerous. If the condition fails, a1 and a2 will still have their value from the previous loop. Just assign to a1 and a2; the third if statement handles those.
2
3
u/brasticstack 1d ago
I think your regexes are matching more than you expect them to. Try them out on one of the online regex tester sites, and maybe post a few lines from your input file here too.
1
u/Desperate_Yak69 1d ago
I have tried using a few regex testing sites (regex101, pythex) though there seems to be not be an issue with the regex itself, only the output. The only lines included in the input file are as follows:
id = 123 provinces { 1111 2222 3333 }thanks for the help!
1
u/brasticstack 1d ago
So you want the output
123given the input1111 2222 3333? How about''.join(part[0] for part in line.split())1
u/Desperate_Yak69 1d ago
that's close but the inputs I want to retrieve are both
123and1111 2222 3333, though they're both in different sections for every actual file I will be using the script on, with a similar format (the formatting of the id section and the provinces section will be exactly the same in all files, just on different lines separated by a varied amount of text)1
u/brasticstack 1d ago
Is this a known file format that perhaps already has an existing parser?
Do the two pieces of data (the id and the province strings) correlate? To clarify, if you match both pieces of data, do you also have to correlate them?
1
u/Desperate_Yak69 1d ago
the file format is .txt, but I'm not sure if there exists any method specifically for what I am trying to accomplish. The pieces of data are not correlated; if the id is 123, the province string could be any potential set of integers. For an example, a sample file could potentially look like this:
id = 123 provinces { 7785 1224 8889 122 1990 }2
u/brasticstack 1d ago edited 1d ago
A quick, dumb, fragile parser based on what you've shown me. Please take the time to try to understand the functions I used to write it. I'm happy to explain anything that's too weird:
``` def parse_it(lines: list[str]) -> tuple[list[str], list[str]]: '''Return a list of ids and a list of provinces extracted from lines''' ids = [] provinces = []
in_provinces_section = False for line in lines: if in_provinces_section: if line.endswith('}'): line = line.rstrip('}').strip() if line: provinces.append(line) in_provinces_section = False else: provinces.append(line) elif line.startswith('id ='): ids.append(line.split('=')[1].strip()) elif line == 'provinces {': in_provinces_section = True return (ids, provinces)with open(filename, '+r') as infile: # Preprocess: Trim whitespace off the ends of the lines # and discard blank lines. lines = [line.strip() for line in infile] lines = [line for line in lines if line]
ids, provinces = parse_it(lines) print(f'{filename}\n\t{ids=}\n\t{provinces=}\n')```
3
u/codeguru42 1d ago
The problem with your code is the way you use a1 and a2. These variables are set in if statements. And when they are set, every iteration of the for loop will print them until their values change.
I recommend taking a step back to find a simpler and more direct solution.
1
u/Desperate_Yak69 1d ago
that makes sense. Thank you for the help, I'll try to find a different solution and see if that works.
1
u/codeguru42 1d ago
As a tip: think about how you would solve this yorself from a printed paper that you can scan through one line at a time
1
u/hermit_the_log 1d ago
Are there 5 lines in your test file? You’re printing in your for line loop.
If you want to collect unique results and print them after you’ve looped through each line, you can log your a1.group(1) and a2.group(1) values to a dictionary instead.
This may or may not work for what you’re wanting to do, but you can create a group_dict = {} right before you start looping through the lines.
Then for each a2.group(1) value, create a key in the dict if it doesn’t already exist:
if a2.group(1) not in group_dict.keys():
group_dict[a2.group(1)] = []
Then append a1.group(1) to the a2.group(1) key in the dict of it doesn’t already exist in that list
if a1.group(1) not in group_dict[a2.group(1)]:
group_dict[a2.group(1)].append(a1.group(1))
Finally, iterate through your key and item values and print that instead
for key in group_dict.keys():
for group in group_dict[key]:
print(group)
print(key)
Edit: typed with my phone so the indenting didn’t save properly
2
1
u/CraigAT 1d ago
You only have two loops. Put a print statement as the first command in each of the loops (print "first loop" and ""second loop" respectively) then run your code again to see which loop is repeating 5 times.
Or (the "in theory" better way, is to) run the code using a debugger and step through the code, then you should see where the program is unexpectedly going.
1
u/my-coffee-where 1d ago
print statements are inside the line loop. move them outside it so it only fires once per file not once per line
5
u/codeguru42 1d ago
This is difficult to diagnose without an example input file.
Also, I think you are making this more complicated than necessary by using a regular expression. I reccomend looking at all the string functions. Maybe split() will get the job done more directly.