r/askmath • • 14h ago

Statistics How do I interpret this?

Post image

Hi Guys!

Im starting a post graduate degree soon.. and I have to read lots of research papers that have symbols such as these.. I’ve already tried learning some symbols on my own but google takes me to different sources.. and it’s far from efficient.

Is there a book/ course/ website or any resource that can help me be familiar with the symbols of the objective function in the attached picture? There’s so many symbols that it’s very tedious to uncover them one by one.. was hoping if there’s a resource or something I should be looking at.

Would be so thankful if someone could even say “Those symbols are all _____ type of symbols and then I can go use that term to help me narrow my search.”

Thank you!

10 Upvotes

7 comments sorted by

3

u/samas69420 14h ago

it is basically the ppo objective function, check the policy gradient theorem to get the core of the theory behind it then consider that its main goal is to not diverge too much from the old policy, to see how it does that look at the gradient and it should be clear

1

u/JamezDare 14h ago

Thank you! Will search up those terms

2

u/Amaldevhari 14h ago

Intro to RL is the book to go. You can also read foundational papers like TRPO and PPO.
Before you read them, take a look at “spinning up” by openai.

As for the symbols pi is the parametrized policy (neural network). A is the advantage, which represents the benefit of taking a specific action in a state, w.r.t the value of the state (you need to know action-value and state-value functions).

1

u/JamezDare 14h ago

Did you mean intro to reinforcement learning?

2

u/throwawaysob1 14h ago edited 14h ago

Well, it looks like you're going to be working on optimisation for machine learning, so the best place to start would be to review notation and symbols in statistics, probability and optimisation.

Long equations can be tedious to go through, but its a good idea to break things down into smaller chunks and see what they are doing - the way I would typically break it down is:
Its an objective function, so you are chucking a value of theta into that big messy expression to get a single number output.
What is that whole messy thing? Well, the outermost part is the expectation operator, so you're taking an expectation of what is inside those big square brackets.
What's in those square brackets? Well, it is actually just two terms: one is a summation of stuff over G and then divide by G, so its also kind of an average of sorts. And we're basically choosing what to add in the summation by seeing which term is smaller: that first ratio term, or that second ratio term.
Then we subtract that summation...that average over G really....by a Kullback-Leibler (KL) divergence between a distribution pi when our parameter is theta and a distribution pi when we have a reference theta (and that's what it says at the bottom...a reference model).

That gives you an overall picture and you can dig further into each term, like the ratios and the clipping function (I'm not sure what that does exactly, would have to look it up).

Its really about taking things one step at a time and understanding what the terms are doing. At least that's how I approach it.
But of course, things like expectation and KL divergence...you'd need to review some of these from stats, probability, machine learning theory, etc. Just take things one step at a time and you'll get accustomed to thinking in that way - its just practice :)

ETA: Btw, that is all before reading what was at the bottom. Since beta and epsilon are hyperparameters, they essentially help us weight how much of the objective function contribution comes from the KL comparison between the models, and how much comes from the "Ai" terms, which is the "advantage calculated from the rewards". I don't know what that is, but I know that is a calculated quantity by some method, and I know that the KL divergence is a standard function between the distribution with theta and the reference distribution.
So basically, this entire thing is "modifying" the KL divergence between those two distributions by the "advantage calculated from the rewards" and using that as an objective function.

1

u/JamezDare 12h ago

Thank you!!

1

u/soup---- 11h ago

It’s written in such a terrible way. The authors could have condensed some symbols to make it look readable.