There are various versions of approximation to identity.

Definition

\(\delta_0\) identity is a functional that \(\delta_0(f) = f(0)\) for \(f \in L^1\).

Approximation to the identity is the weak* approximation(?from the perspective of distribution theory

This theorem (technique) enables us to prove the theorem on some smooth functions (nice properties).

Technically, we have a steps to do it.

  1. Find the suitable kernel \({K_{\delta}}\) to approximate identity
  2. convolve the function \(f\) with kernel, thai it define \(f_{\delta}(x):=f*K_{\delta}(x)\)
  3. Prove the desired result on a smooth function
  4. take the limit and see whether the property desired is preserved.

We want our kernel to have these properties and then they can perform as the identity while taking limit.

well we first define a function \(K_\delta\) satisfies

  1. \(\displaystyle \int_{\mathbb{R}^n} K_{\delta}\,\mathrm{d}x = 1\);
  2. \(\sup \lVert K_{\delta} \rVert_{1} < \infty\) If \(K_{\delta}\) is non-negative, then this is not necessary.
  3. \(\displaystyle \int_{\lvert x\rvert>\eta} \lvert K_\delta\rvert\,\mathrm{d}x \to 0\) as \(\delta \to 0\).

This results directly:

Theorem

For any bounded continuous function \(\phi(x)\), \(\phi*K_{\delta}(x) \rightarrow \phi(x)\) as \(\delta \rightarrow 0\) uniformly.

The proof of it is easy. One can show it by seperating the convolution into two parts. Given \(\epsilon >0\),

  1. the part that dominated by the decline of \(K_{\delta}\)
  2. the part that dominated by the continuity of \(f(x)\)

E.g \({(n+1)x^n}\) supported on \([0,1]\)

So one can say: \(\int_{0}^{1} f(1-x)(n+1)x^n \,\mathrm{d}x \rightarrow f(1)\)

In Stein’s book, One can find a noneagtive function \(\phi(x)\) (may not be a function, it can be a measure in the functional form) supported in unit ball \(B(0,1)\) such that

  1. \(\int_{\mathbb{R}^n} \phi(x) \,\mathrm{d}x = 1\)

Here, if we take the integral equals A>0, we may approach a

Then one may define \(\phi_{\epsilon}(x) = \epsilon^{-n}\phi(x/\epsilon)\), this is a dilation of \(\phi(x)\), this dilation satisfies the above 3 conditions.

\[ \hat\phi_\epsilon(x) = \hat\phi(\epsilon x) \]

let \(\epsilon \rightarrow 0\), then \(\hat\phi_\epsilon(x) \rightarrow \hat\phi(0) = \int \phi \,\mathrm{d}x = 1\), (Here we use the continuity of the Fourier transformed function of \(L^1\)) this has the same Fourier transformation as Derac measure. By the uniqueness of Fourier transformation

of meaure in \(M(\mathbb{R}^n)\),

\[ \hat\mu = \hat\nu \Rightarrow \mu = \nu \]

we have \(\phi_\epsilon(x)\) converges weakly to \(\delta_0\).

The above technique is applied to see the distribution of the sum two independent random varibale is a the convolution of each distribution.

Theorem

Suppose that \(X_{1},X_{2}\) are two independent random variable, then the distribution of \(X_{1}+X_{2}\) satisfies:

\[ P_{X_{1}+X_{2}} = P_{X_{1}}*P_{X_{2}} \]

Proof: Consider the Fourier transformtion of \(\widehat{P_{X_{1}+X_{2}}}\), and

\[ \widehat{P_{X_{1}+X_{2}}} =\widehat{P_{X_{1}}} \widehat{P_{X_{2}}} = =\widehat{ P_{X_{1}}*P_{X_{2}}} \]

The first equality relies on the independence and the second equality relies on the product rules of the Fourier transformation of measure in \(M(\mathbb{R}^n)\). Then similarily by the uniqueness, we obtain the desired result.

Okay, now we state another thorem of approximation to identity. This enables us to do the same thing for functions in \(L^1\) with few adjustment.

  1. \(\lvert K_{\delta}(x)\rvert\leq A\delta^{-d}\) for all \(\delta >0\)

  2. \(\lvert K_{\delta}(x)\rvert\leq \dfrac{A\delta}{\lvert x\rvert^{d+1}}\) for all \(\delta >0\) and \(x\in \mathbb{R}^d\)

This adjustment is also a classical technique when we deal with the estimation problem, applying two inequalities with the same direction and use them to estimate different parts of the problem.

One can check that this adjustment still implies the condition 2 and 3 in previous theorem.

approximation for \(L^1\) functions (THM 2.1 on Stein’s Real analysis)

If \({K_{\delta}}\) is an approximation to the identity and \(f\) is integrable on \(\mathbb{R}^d\), then

\[ (f*K_{\delta}(x))\rightarrow f(x) \quad\text{ as } \delta \rightarrow 0 \]

for every x in the Lebesgue set of \(f\). Since \(a.e \,x\) is in the Lebesgue set, the limit holds for \(a.e. \,x\)

The idea of the proof inherits from above, esimate separately. (The proof of DCT is similar and The proof of the inverse formula of Fourier trasnformation is also similar)

We will esimate the difference near \(x\) and way far from \(x\).

proof: \(\lvert f*K_{\delta}-f\rvert\le \int \lvert f(y-x)-f(x)\rvert\,\lvert K_{\delta}(x)\rvert\,\mathrm{d}x\)

Recall that we have

  1. \(\lvert K_{\delta}(x)\rvert\leq A\delta^{-d}\) for all \(\delta >0\)

  2. \(\lvert K_{\delta}(x)\rvert\leq \dfrac{A\delta}{\lvert x\rvert^{d+1}}\) for all \(\delta >0\) and \(x\in \mathbb{R}^d\)

Since this, we try to esimate the right handside of the inequality by separating into ball and dyacdic annuluses.

\[ \int \lvert f(y-x)-f(x)\rvert\,\lvert K_{\delta}(x)\rvert\,\mathrm{d}x = \int_{\lvert x\rvert\le\delta}\lvert f(y-x)-f(x)\rvert\,\lvert K_{\delta}(x)\rvert\,\mathrm{d}x+\sum_{j>1}\int_{2^{j-1}\delta<\lvert x\rvert\le2^{j}\delta}\lvert f(y-x)-f(x)\rvert\,\lvert K_{\delta}(x)\rvert\,\mathrm{d}x \]

The first term:

\(\lesssim \dfrac{1}{\delta^d}\int_{\lvert x\rvert\le\delta}\lvert f(y-x)-f(x)\rvert\,\mathrm{d}x\)

The second term: We apply 3.

\(\lesssim \sum \dfrac{1}{{2^{j(d+1)}\delta}^{d}}\int_{\lvert x\rvert\le 2^{j}\delta}\lvert f(y-x)-f(x)\rvert\,\mathrm{d}x\)

Obstabcle we have now or we desire to show that \(\dfrac{1}{\delta^d}\int_{\lvert x\rvert\le\delta}\lvert f(y-x)-f(x)\rvert\,\mathrm{d}x\)

We can further our disscusion over mollifers, these are functions that allow us to improve the property of less smooth functions.

For example:

\[ \eta(x):= \exp\left(\frac{1}{\lvert x\rvert^2-1}\right) \quad \text{when } x \in [-1,1] \]

Let \(C\) be the integral of \(\eta(x)\) over \(I\) (\(C \approx 0.4440\)), we can always normalize the \(\eta(x)\)