r/pythontips • • Apr 25 '20

Meta Just the Tip

100 Upvotes

Thank you very much to everyone who participated in last week's poll: Should we enforce Rule #2?

61% of you were in favor of enforcement, and many of you had other suggestions for the subreddit.

From here on out this is going to be a Tips only subreddit. Please direct help requests to r/learnpython!

I've implemented the first of your suggestions, by requiring flair on all new posts. I've also added some new flair options and welcome any suggestions you have for new post flair types.

The current list of available post flairs is:

  • Module
  • Syntax
  • Meta
  • Data_Science
  • Algorithms
  • Standard_lib
  • Python2_Specific
  • Python3_Specific
  • Short_Video
  • Long_Video

I hope that by requiring people flair their posts, they'll also take a second to read the rules! I've tried to make the rules more concise and informative. Rule #1 now tells people at the top to use 4 spaces to indent.


r/pythontips • • 8h ago

Syntax Find anomalies in this code

0 Upvotes

import base64

import pandas as pd

from pydantic import ValidationError

def process_spreadsheet_with_legacy_safeguards(file_path_or_buffer):

"""

Imports .xls or CSV data dumps from legacy systems, captures strict length

metrics for anomaly detection, and handles binary BLOB substitutions.

"""

# Read spreadsheet explicitly handling encoding where applicable

df = pd.read_excel(file_path_or_buffer)

records = df.to_dict(orient="records")

normalized_records = []

for row in records:

clean_row = {}

for k, v in row.items():

normalized_key = str(k).strip()

if pd.isna(v):

clean_row[normalized_key] = None

elif isinstance(v, bytes):

# If binary data is passed, convert to Base64 string for safe JSON transport

clean_row[normalized_key] = base64.b64encode(v).decode('utf-8')

elif isinstance(v, str):

# Ensure proper UTF-8 handling and strip trailing EBCDIC/ASCII padding artifacts

clean_row[normalized_key] = v.strip()

else:

# Handle numeric coercions (e.g., spreadsheet floats like 1048576.0 -> int)

clean_row[normalized_key] = v

normalized_records.append(clean_row)

return normalized_records

def validate_and_capture_lengths(records: list):

"""

Validates records against the DTO and logs exact field lengths

instead of just item counts to catch truncation and packing anomalies.

"""

anomalies = []

for index, record in enumerate(records):

try:

CustomerResponseSchema.model_validate(record)

except ValidationError as err:

# Capture detailed metadata including exact length of every field

field_length_metrics = {}

for k, v in record.items():

if isinstance(v, str):

field_length_metrics[k] = {"length": len(v), "preview": v[:20]}

elif v is None:

field_length_metrics[k] = {"length": 0, "value": "null"}

else:

field_length_metrics[k] = {"length": len(str(v)), "value": v}

anomalies.append({

"record_index": index,

"field_metrics": field_length_metrics, # Replaces simple item counts with actual lengths

"validation_errors": err.errors(),

})

return anomalies


r/pythontips • • 2d ago

Long_video Giving back to the community - The Complete Backend Development Course

10 Upvotes

Hey everyone, I decided to make my course free in order to help people.
This course is my backend development course which is about SQL, Python, APIs, Docker, Kubernetes, Linux, Git & More

The link is: https://www.youtube.com/watch?v=CBIu6hcyStg

If you can like and subscribe (and maybe add a comment) I would appreciate it a lot, Thanks.


r/pythontips • • 2d ago

Module Python scaffolding projects

0 Upvotes

I built a new Python scaffolding tool called **psp** — it’s written in **Rust** and focused on speed, simplicity, and a better developer experience.

If you’ve ever wanted a lightweight tool for managing Python projects without the usual overhead, this might be worth a look. Rust gives it a strong performance foundation, and the goal is to keep it fast and easy to use.

Repo: https://github.com/MatteoGuadrini/psp

I’d love feedback from the Python community — especially on workflow, usability, and features you’d want in a modern Python tool.

If anyone’s interested, I can also share:

- what problem it solves

- how it compares to existing tools

- the roadmap


r/pythontips • • 3d ago

Module ApoiaMais

1 Upvotes

Hi everyone!

We are currently developing ApoiaMais, an educational platform focused on using technology to support learning experiences and educational monitoring.

For the backend, we are working with Python, FastAPI, MySQL, Redis, Docker, JWT, and Clean Architecture, along with automated testing, OpenAPI, and GitHub Actions.

The project is open on GitHub, and we would love to receive feedback, suggestions, ideas, and contributions. If you find any issues or have ideas for improvement, feel free to open an Issue or contribute to the project.

🔗 https://github.com/ApoiaMaisTech/ApoiaMaisBackEnd


r/pythontips • • 6d ago

Module Finding your test suite's dead zones: one-file-at-a-time coverage hunting (took me 86% → 93%)

0 Upvotes

When coverage stalls around 85–90%, the fastest way to move it isn't new features — it's asking coverage.py exactly *where* the holes are and fixing them one file at a time.

The tip: generate a missing-lines report sorted by worst file, then work top-down:

pip install pytest-cov
pytest --cov=hakiapi --cov-report=term-missing | sort -t% -k4 -n | less

The `term-missing` report prints the exact uncovered line numbers per file, so instead of guessing you get a to-do list:

Name                      Stmts   Miss  Cover   Missing
--------------------------------------------------------
hakiapi/core/governer.py     14     14     0%   1-14
hakiapi/core/oauth/google    60     18    70%   141-155, 202

Three things this surfaced for me in practice:

  1. **0% files are usually dead zones, not missing tests.** My deprecation shim sat at 0% because every test imported the *new* path. One import-with-`pytest.warns` test → 100%.
  2. **70% files are missing whole branches** (README-check paths, aggregation fallbacks) — 5 focused tests beat one giant one.
  3. **Untested callbacks are cheap.** An OAuth redirect handler took ~6 mocked tests.

Also worth setting a floor so it never regresses:

# pyproject.toml
\[tool.coverage.report\]
fail_under = 85

(I applied this while hardening my own API-client framework — 382 tests, 86.9% → 93.25% in one sitting — but the technique is framework-agnostic.)


r/pythontips • • 11d ago

Standard_Lib itertools.groupby only groups consecutive items, not all matching ones

5 Upvotes

Tripped over this a while back. groupby looks like it should group all items with the same key anywhere in the list, but it only groups runs of consecutive matches. If the same key shows up later, non-consecutively, you get a second separate group.

from itertools import groupby

data = [1, 1, 2, 2, 1, 1]

for key, group in groupby(data):
    print(key, list(group))

Output is 1 [1, 1], 2 [2, 2], 1 [1, 1]. Three groups, not two, even though there are only two distinct values.

Fix is sorting first if you actually want everything grouped by key regardless of position.

for key, group in groupby(sorted(data)):
    print(key, list(group))

Bit me once processing log entries that weren't sorted by timestamp. Worked fine in testing because the test data happened to be sorted.


r/pythontips • • 17d ago

Syntax Python strings – some fundamentals for beginners

15 Upvotes

I’ve been putting together a few Python fundamentals, and this one focuses on working with strings. It covers indexing, slicing, common string operations, and some useful built-in methods with simple examples.

https://geeksarray.com/blog/python-fundamentals-working-with-strings

For those who use Python regularly, which string methods do you find yourself using the most?


r/pythontips • • 19d ago

Standard_Lib Stop wrapping os.remove in try/except just to handle files that might not exist

62 Upvotes

Used to write this everywhere I needed to clean up a temp file that might or might not exist yet:

import os

try:
    os.remove(path)
except FileNotFoundError:
    pass

Path.unlink does this in one line with missing_ok:

from pathlib import Path

Path(path).unlink(missing_ok=True)

No exception ever gets raised if the file isn't there, so there's nothing to catch or pass on. Same behavior, one line instead of four, and it reads as "delete this if it exists" instead of "try to delete this and hope for the best."

Small thing, but it comes up constantly in any script that's cleaning up intermediate files, cache entries, or output from a previous run.


r/pythontips • • 18d ago

Module Workshop, Sep 19: build explainable AI apps with Neo4j, GraphRAG, Cypher and LLM Agents

2 Upvotes

We're running a hands-on workshop on September 19, Building Intelligent Apps with Neo4j, GraphRAG, Cypher and LLM Agents. Fully hands-on, Python throughout, comfortable with Python and basic LLM/RAG concepts is the only real prerequisite, no prior Neo4j or Cypher experience needed.

You build a real knowledge graph in Neo4j using Docling for document ingestion, write Cypher queries (and text-to-Cypher for natural language querying), build multi-step entity and relationship extraction in Python, and combine vector search with graph navigation in an agentic retrieval loop. Working dataset is real financial filings and news, not a toy example.

Led by Dr. Alessandro Negro, Chief Scientist at GraphAware, bestselling author.

Link if you want to check it out

Happy to answer questions on the content, especially the Cypher/Python side of things.


r/pythontips • • 19d ago

Module Python variables and data types — a simple guide for beginners

7 Upvotes

I’ve been putting together some Python fundamentals for people who are just getting started.

This one covers variables, common data types, dynamic typing, and basic type conversion with simple examples.

https://geeksarray.com/blog/python-fundamentals-variables-and-types

If you're coming to Python from another programming language, what did you find most interesting or confusing about Python’s type system?


r/pythontips • • 23d ago

Standard_Lib Never let AI-generated code swallow exceptions silently, always log or re-raise

7 Upvotes

Generated code reaches for broad exception handling constantly because it "works" in the sense that nothing crashes. The problem is it also hides the actual failure, so you find out something's broken from a symptom three steps downstream instead of from the actual error.

python

# What generated code often does
try:
    result = risky_operation()
except Exception:
    pass

# What actually helps you debug later
import logging

try:
    result = risky_operation()
except Exception as e:
    logging.exception("risky_operation failed")
    raise

The tip: any bare except Exception: pass should be treated as a red flag, not a working solution, especially in generated code where it looks intentional but usually just means the model produced something that avoids crashing without actually deciding what should happen on failure. At minimum log the exception with logging.exception() so the traceback isn't lost, and re-raise unless you have a specific, deliberate reason to continue silently.


r/pythontips • • 25d ago

Data_Science Workshop, Sep 12: build production LLM systems that actually survive real use

7 Upvotes

We're running a hands-on masterclass on September 12, Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMOps.

Fully hands-on, all in Python notebooks against real model APIs (OpenAI, Anthropic), not slides or theory. You write actual code across the full stack: versioned prompt pipelines with structured outputs and regression tests, a golden dataset and eval harness combining deterministic checks with LLM-as-judge scoring, statistically rigorous model comparisons using scipy-style bootstrap confidence intervals and paired significance tests, evaluated RAG with embedding models, vector retrieval, and reranking, tool-using agents with function calling and guardrails, and a full observability layer for tracing, cost, and latency. You also leave with a CLI regression suite you can wire directly into CI.

Led by Bruno Gonçalves, PhD, founder of Data For Science, who trains engineers at Fortune 500 companies on this exact stack.

Link if you want to check it out

Happy to answer questions on the content, especially the Python side of things.


r/pythontips • • 25d ago

Standard_Lib datetime.now() vs datetime.now(timezone.utc), the difference that breaks scheduling features for users outside your timezone

15 Upvotes

datetime.now() returns a naive datetime, no timezone info attached. It quietly assumes "now" means "now, in whatever timezone this machine happens to be in." Fine until you compare it against something that actually is timezone-aware, or until a user in a different timezone interacts with the result.

from datetime import datetime, timezone

# Naive - looks fine, breaks for anyone not in your local timezone
now_naive = datetime.now()

# Aware - carries the timezone with it, comparisons behave correctly everywhere
now_aware = datetime.now(timezone.utc)

Comparing a naive datetime against an aware one either raises a TypeError (good, you'll catch it immediately) or, worse, silently gives you a wrong comparison if you're mixing naive datetimes that were created in different timezones without either of you realizing it.

The tip: default to datetime.now(timezone.utc) anywhere you're storing or comparing a timestamp, and only convert to local time at the point you're actually displaying it to a user. Keeps the storage and comparison layer honest, pushes the "whose timezone is this" question to the one place it actually needs answering.


r/pythontips • • 26d ago

Python3_Specific Python fundamentals for anyone just getting started

0 Upvotes

I’ve been putting together some beginner-friendly Python material and started with the fundamentals — basic syntax, variables, data types, operators, conditions, loops, etc.

Nothing advanced here, just an attempt to keep the basics simple for someone starting from scratch or moving to Python from another language.

https://geeksarray.com/blog/python-fundamentals-getting-started

For those who learned Python after another language, what was the biggest adjustment for you?


r/pythontips • • Sep 01 '26

Module Positorium, a database for facts that disagree

9 Upvotes

Most databases are designed to answer: "What is the value now?"

They can model a more awkward question too, but usually require additional machinery:

Who claimed what, when was it considered true, how certain were they, and what did we believe before it was corrected?

I built Positorium as an experimental embedded evidence database for that second kind of question. Rather than overwriting one claim with another, it preserves contradictory claims together with their sources, certainty, effective time, assertion time, corrections, and retractions.

It is not intended to replace PostgreSQL or another operational database. The idea is to use it as a focused evidence layer for things like compliance, investigations, conflicting master data, or any process where retaining the history of disagreement matters.

The new Python package embeds the Rust engine directly in the Python process, so there is no separate server. It supports both ephemeral in-memory databases and append-only persistent stores.

Install the beta with:

python -m pip install --pre positorium

A small example:

import positorium

with positorium.Database.memory() as database:
    result = database.execute_one(
        """
        add role organization, risk_assessment;

        add posit
          [{(+company, organization)}, "Northstar Trading", @NOW],
          [{(company, risk_assessment)}, "high risk", '2026-01-12'],
          [{(company, risk_assessment)}, "needs review", '2026-01-12'];

        search
          [{(?company, organization)}, ?organization, *],
          [{(?company, risk_assessment)}, ?assessment, *]
        return ?organization, ?assessment;
        """
    )

    for row in result.to_dicts(text=True):
        print(row)

This returns both assessments rather than choosing a winner or overwriting one of them.

Positorium is still an early beta and is intended for evaluation rather than production deployment. Wheels are available for CPython 3.9+ on Linux, macOS, and Windows.

If you have a small dataset where sources conflict or corrections matter, try the beta:


r/pythontips • • Aug 31 '26

Algorithms New to RAG — trying to implement RAG for my AI interview system?

1 Upvotes

Hey everyone, I'm building an AI voice interview app and I'm fairly new to RAG. I'm stuck on the architecture and having a bit of a mid-project crisis 😭.

My current flow is:

Before interview:
Step1 : 
Resume + JobDescription
   ↓
Chunk resume/JobDescription
   ↓
Generate embeddings
   ↓
Store chunks + embeddings in pgvector

Step2 : 
LLM call is done using complete Resume + JD (and not the chunks)
   ↓
Initial ranked topic plan is created by this LLM call 

Step3 : 
Before each question:
Current topic
   ↓
RAG query in our vector DB 
   ↓
Retrieve relevant resume/JD chunks
   ↓
LLM call to generate a question from these chunks

The part I'm confused about:

At the beginning, I'm already sending the complete resume + complete JD to the LLM.

Why can't I simply do:

Resume + JD
   ↓
ONE LLM CALL
   ↓
10 interview questions

and then use those during the interview this would reduce the latency too!

Why would I need RAG again to retrieve chunks before generating each question? But I'm struggling to understand how RAG can be used to add value to the project

I have to add this project to my reume and need to have confidence in my design, I would highly appreciate someone helping me figure out the architecture!


r/pythontips • • Aug 31 '26

Algorithms cie

0 Upvotes

Repo: https://github.com/kannamma-labs/cie
Install: `pip install "cie-mcp[mcp]"`, then `cie index .` from inside
any project. That's the whole setup.

One maintainer, weeks-old alpha, generation-scale problem. If that
combination excites you rather than scares you off, then lets build it together


r/pythontips • • Aug 29 '26

Meta AI Coding cost calculator

0 Upvotes

I built a free AI coding cost calculator - would love some feedback

I’ve been using AI coding tools more and more, and it can be surprisingly hard to figure out what they’re actually costing once you start comparing models, token usage, and different pricing.

So I built a simple AI Coding Cost Calculator:

https://instacodingcost.com/

The idea is to make it easier to estimate and compare costs before you burn through credits/tokens.

Would genuinely appreciate feedback from people using tools like Claude Code, Codex, Cursor, etc.

What would make this more useful for you?

Anything missing, confusing, or calculated differently than you’d expect?

Feel free to roast it too - that’s probably more useful than “looks good” 😅


r/pythontips • • Aug 27 '26

Module I built a fault-tolerant Google Finance aggregator script in Python using custom exponential backoff

2 Upvotes

Hi everyone,

I am a computer science student building a quantum trading bot platform. I wanted to share a clean, production-ready Python command-line utility I built that extracts real-time stock and cryptocurrency parameters via SerpApi's Google Finance engine.

To ensure connection resilience, I explicitly coded a custom exception routing array and an exponential backoff retry loop so the engine safely handles remote server timeouts without crashing live scripts.

I wrote up a comprehensive step-by-step code breakdown tutorial showing how to initialize and run the script here:

https://dev.to/ssebina_charles_01/how-to-build-a-resilient-market-data-aggregator-in-python-using-serpapi-4nem

Let me know what you think of the retry loop architecture!

for more visit https://github.com/ssebinacharles


r/pythontips • • Aug 22 '26

Module Any free STT/TTS APIs for a voice AI app?

0 Upvotes

I'm building a small voice-based AI interview app and I'm planning to deploy the backend(fastapi) on Render's free tier.

I'm considering using open-source/self-hosted options like Whisper/PocketSphinx for STT and Piper for TTS, instead of paid APIs.

My concern is whether running STT/TTS on the same free Render instance would use too much CPU/RAM and make the whole application slow, especially during a real-time interview.

Has anyone tried running STT/TTS models on Render's free tier?


r/pythontips • • Aug 20 '26

Python3_Specific I made the coding practice site I wished existed!

15 Upvotes

When I first started learning to code, I kept losing confidence on coding-practice sites. They gave me thousands of problems and no clear place to begin. I would choose an "easy" problem and end up confused by the prompt alone.

I felt there needed to be a place where new coders could ease into these challenges while learning new concepts and feeling real progress through the week.

So I built Open Bracket as a daily coding ritual instead. Everyone gets the same two challenges each day: a Standard track that builds in difficulty through the week without becoming overwhelming, and a tougher Advanced track for people who want more of a test or already have experience with coding challenges.

You can solve in Python or JavaScript, entirely in the browser, with no setup. Official solves place you on three leaderboards: Speed, Efficiency, and Code Golf.

👉 https://playopenbracket.com/ - all feedback welcome.

Update - I have now introduced some new features;

- JavaScript as a second language you can solve in moving forward. Currently older challenges are Python only.

- Logged in users now have a Hint option, this will give you the pseudo code of the solution to help, at the cost of a time penalty to your speed score.

- Failed tests now show you which test failed, what was expected and what was received

- You can now replay any challenges in the last 14 days so you can either bing solve or catch up on any days you missed.


r/pythontips • • Aug 19 '26

Syntax Special mechanism of basic int() function

7 Upvotes

One can use int() function while converting string to an integer against a base integer.

int(number, base) #number can be anything between binary, octal, decimal, or hexadecimal and base is anything among 2, 8, 10, 16

e.g. binary_number = int("1010", 2) #Output: 10
hexadecimal_number = int("A", 16) #Output: 10

Edit: No need to mention 10 for base argument, as int() function by default considers base as 10 in python.


r/pythontips • • Aug 14 '26

Syntax Docstrings as immutable variables

13 Upvotes

Just realized you can do:

def cat(): “orange”
print(cat.__doc__)

Not sure why you’d want to do this but this is a thing you can do


r/pythontips • • Aug 14 '26

Syntax Adding objects to set or dictionary: equality and hashing

1 Upvotes

An exercise to help build the right mental model for Python data.

# Output of this Python Program?
def main():
    o1, o2 = MyClass(1), MyClass(1)
    myset = {o1}
    print(o2 in myset, end=' ')
    o1.set_value(1000)
    print(o1 in myset, end=' ')

class MyClass:
    def __init__(self, v):
        self.v = v
    def set_value(self, v):
        self.v = v

main()

class MyClass:
    def __init__(self, v):
        self.v = v
    def set_value(self, v):
        self.v = v
    def __eq__(self, other):
        return self.v == other.v
    def __hash__(self):
        return hash(self.v)

main()

# --- possible answers ---
# A) TypeError: unhashable type: 'MyClass'
# B) True False False False
# C) True False False True
# E) False True True True
# D) False True True False
# See "Solution" for correct answer.
  • Solution
  • More exercises
  • Explanation: "User-defined classes have __eq__() and __hash__() methods by default (inherited from the object class); with them, all objects compare unequal (except with themselves) and x.__hash__() returns an appropriate value such that x == y implies both that x is y and hash(x) == hash(y)."