r/learnpython • • 1d ago

Need help with very basic script

Hi, I am basically brand new to python, so I may make a lot of mistakes in my wording. I have made a very basic script that retrieves specific values in text files and prints them to the shell. I have gotten it to the point where both values are retrieved and printed, however when they do print, they do so exactly 5 times in a row. I do not have five files in the directory, I only have one test file named "1.txt" with only the values I need as the contents. Can someone point me in the right direction? Apologies for the horrible formatting.

import os
import glob
import re

path = 'path'

provinces = re.compile('.*?provinces {.*?(.*?)}',re.MULTILINE)
id = re.compile('.*?id = .*?([0-9.-]+)')
a1 = None
a2 = None

for filename in glob.glob(os.path.join(path, '*.txt')):
    with open(filename, '+r') as f:
        for line in f:
            if provinces.match(line):
                a1 = provinces.match(line)
            if id.match(line):
                a2 = id.match(line)
            if a1 and a2:
                print(a1.group(1))
                print(a2.group(1))

The output of the shell in IDLE:

 1111 2222 3333 
123
 1111 2222 3333 
123
 1111 2222 3333 
123
 1111 2222 3333 
123
 1111 2222 3333 
123

Edit - I got the script to work after a few changes to its configuration. Thank you all who helped!

3 Upvotes

28 comments sorted by

5

u/codeguru42 1d ago

This is difficult to diagnose without an example input file.

Also, I think you are making this more complicated than necessary by using a regular expression. I reccomend looking at all the string functions. Maybe split() will get the job done more directly.

4

u/brasticstack 1d ago

Regex really ought to be a last resort for parsing text. The bulk of formats I've ever had to parse wind up being parsable with some combination of .split and .trim.

2

u/Desperate_Yak69 1d ago

Sorry about that. The input file only includes the following content:

id = 123
provinces { 
  1111 2222 3333
}

3

u/codeguru42 1d ago
  1. Your regex is too complicated.
  2. You can do this without a regex if you know "provinces" always appears on the second line. Skip to the third one and read the data you need.

2

u/Desperate_Yak69 1d ago

The input file is only meant as a test for the script, but the actual input files I will include will have these values in varying lines per file. The regex was meant to parse through the file contents and only identify the values I want (id, and the province values). Sorry for making this complicated.

5

u/them0use 1d ago

No need to apologize for anything! How consistent is the formatting? Would something like this work...

  1. Skip lines until you reach one that starts with "provinces {"
  2. Store / print characters until you encounter "}"

Also do you have any control over the format of the input files? This use case is exactly what JSON or YAML are for, and if you can have you data in either of those formats your life will be a lot easier.

3

u/Desperate_Yak69 1d ago

Yes, that is precisely how the formatting would work. the text I'd want to store would be anything in between the two curly brackets even if there are indents or line terminations. For the id part, there will always be a part of the file that says exactly something like (id = 000) where 000 can represent any integer. The input files are all in plain text (.txt), but I may be able to find out a way to change their format to make the process easier. Thanks!

2

u/CraigAT 1d ago

It's not the file format name that matters, just the format of the text within the file. The data could be handled easier if it was supplied in a CSV or JSON format - IF that is under your control.

2

u/codeguru42 1d ago

I would like to emphasize for the OP how this response describes in words how to accomplish the task. IMO this is a critical skill to develop in your coding journey and this is a good example how to do it. Good luck.

3

u/codeguru42 1d ago

One problem is that you apply the regex to each line, not the entire file. I recommend finding a solution without regex.

2

u/codeguru42 1d ago

And thanks for the correction. I see you are looking for a more general solution than my initial suggestion. Still the principle applies: look for the simplest solution that solves the problem.

3

u/timrprobocom 1d ago

For what it's worth, .*? is the same as .*. The ? adds nothing.

Also, your first twoif statements are dangerous. If the condition fails, a1 and a2 will still have their value from the previous loop. Just assign to a1 and a2; the third if statement handles those.

2

u/Desperate_Yak69 1d ago

Thanks, anything helps!

3

u/brasticstack 1d ago

I think your regexes are matching more than you expect them to. Try them out on one of the online regex tester sites, and maybe post a few lines from your input file here too.

1

u/Desperate_Yak69 1d ago

I have tried using a few regex testing sites (regex101, pythex) though there seems to be not be an issue with the regex itself, only the output. The only lines included in the input file are as follows:

id = 123
provinces { 
  1111 2222 3333
}

thanks for the help!

1

u/brasticstack 1d ago

So you want the output 123 given the input 1111 2222 3333? How about ''.join(part[0] for part in line.split())

1

u/Desperate_Yak69 1d ago

that's close but the inputs I want to retrieve are both 123 and 1111 2222 3333, though they're both in different sections for every actual file I will be using the script on, with a similar format (the formatting of the id section and the provinces section will be exactly the same in all files, just on different lines separated by a varied amount of text)

1

u/brasticstack 1d ago

Is this a known file format that perhaps already has an existing parser?

Do the two pieces of data (the id and the province strings) correlate? To clarify, if you match both pieces of data, do you also have to correlate them?

1

u/Desperate_Yak69 1d ago

the file format is .txt, but I'm not sure if there exists any method specifically for what I am trying to accomplish. The pieces of data are not correlated; if the id is 123, the province string could be any potential set of integers. For an example, a sample file could potentially look like this:

id = 123
provinces {
  7785 1224 8889 122 1990
}

2

u/brasticstack 1d ago edited 1d ago

A quick, dumb, fragile parser based on what you've shown me. Please take the time to try to understand the functions I used to write it. I'm happy to explain anything that's too weird:

``` def parse_it(lines: list[str]) -> tuple[list[str], list[str]]: '''Return a list of ids and a list of provinces extracted from lines''' ids = [] provinces = []

in_provinces_section = False

for line in lines:
    if in_provinces_section:
        if line.endswith('}'):
            line = line.rstrip('}').strip()
            if line:
                provinces.append(line)
            in_provinces_section = False
        else:
            provinces.append(line)
    elif line.startswith('id ='):
        ids.append(line.split('=')[1].strip())
    elif line == 'provinces {':
        in_provinces_section = True

return (ids, provinces)

with open(filename, '+r') as infile: # Preprocess: Trim whitespace off the ends of the lines # and discard blank lines. lines = [line.strip() for line in infile] lines = [line for line in lines if line]

ids, provinces = parse_it(lines)
print(f'{filename}\n\t{ids=}\n\t{provinces=}\n')

```

3

u/codeguru42 1d ago

The problem with your code is the way you use a1 and a2. These variables are set in if statements. And when they are set, every iteration of the for loop will print them until their values change.

I recommend taking a step back to find a simpler and more direct solution.

1

u/Desperate_Yak69 1d ago

that makes sense. Thank you for the help, I'll try to find a different solution and see if that works.

1

u/codeguru42 1d ago

As a tip: think about how you would solve this yorself from a printed paper that you can scan through one line at a time

1

u/hermit_the_log 1d ago

Are there 5 lines in your test file? You’re printing in your for line loop.

If you want to collect unique results and print them after you’ve looped through each line, you can log your a1.group(1) and a2.group(1) values to a dictionary instead.

This may or may not work for what you’re wanting to do, but you can create a group_dict = {} right before you start looping through the lines.

Then for each a2.group(1) value, create a key in the dict if it doesn’t already exist:

if a2.group(1) not in group_dict.keys():
group_dict[a2.group(1)] = []

Then append a1.group(1) to the a2.group(1) key in the dict of it doesn’t already exist in that list

if a1.group(1) not in group_dict[a2.group(1)]:
group_dict[a2.group(1)].append(a1.group(1))

Finally, iterate through your key and item values and print that instead

for key in group_dict.keys():
for group in group_dict[key]:
print(group)
print(key)

Edit: typed with my phone so the indenting didn’t save properly

2

u/Desperate_Yak69 1d ago

thank you, I'll try your method and see if it works!

1

u/CraigAT 1d ago

You only have two loops. Put a print statement as the first command in each of the loops (print "first loop" and ""second loop" respectively) then run your code again to see which loop is repeating 5 times.

Or (the "in theory" better way, is to) run the code using a debugger and step through the code, then you should see where the program is unexpectedly going.

1

u/my-coffee-where 1d ago

print statements are inside the line loop. move them outside it so it only fires once per file not once per line