[Text extracted through search] · (III) BERT-Based Text Sampling

Starting with this post, we will begin applying the sampling algorithms introduced earlier to concrete examples of text generation. As our first example, we'll look at how BERT can be used to perform random text sampling. By "random text sampling," we mean randomly generating natural language sentences from a model. The common view is that this kind of random sampling is a capability unique to unidirectional autoregressive language models like GPT2 and GPT3, and that bidirectional masked language models (MLMs) like BERT simply cannot do it.

Is that really true? Of course not. It turns out that BERT's MLM can also be used to sample text — in fact, it amounts to exactly the Gibbs sampling procedure introduced in the previous post. This was first clearly pointed out in the paper BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model. The paper's title is quite amusing too: "BERT also has a mouth, so it must say something." Let's now see what BERT actually has to say~ more

Sampling Procedure

First, let's revisit the Gibbs sampling procedure introduced in the previous post:

Gibbs Sampling
The initial state is $\boldsymbol{x}_0=(x_{0,1},x_{0,2},\cdots,x_{0,l})$, and the state at time $t$ is $\boldsymbol{x}_t=(x_{t,1},x_{t,2},\cdots,x_{t,l})$.
We sample $\boldsymbol{x}_{t+1}$ via the following procedure:
1. Uniformly sample an index $i$ from $1,2,\cdots,l$;
2. Compute $p(y|\boldsymbol{x}_{t,-i})=\frac{p(x_{t,1},\dots,x_{t,i-1},y,x_{t,i+1},\cdots,x_{t,l})}{\sum\limits_y p(x_{t,1},\dots,x_{t,i-1},y,x_{t,i+1},\cdots,x_{t,l})}$;
3. Sample $y\sim p(y|\boldsymbol{x}_{t,-i})$;
4. $\boldsymbol{x}_{t+1} = {\boldsymbol{x}_t}_{[x_{t,i}=y]}$ (i.e., replace the $i$-th position of $\boldsymbol{x}_t$ with $y$ to obtain $\boldsymbol{x}_{t+1}$).

The most crucial step here is computing $p(y|\boldsymbol{x}_{-i})$, which specifically means "predicting the $i$-th element using the remaining $l-1$ elements obtained by removing the $i$-th element." Readers familiar with BERT should immediately recognize this: isn't this exactly what BERT's MLM is designed to do? So combining the MLM with Gibbs sampling indeed lets us perform random sampling of text.

So, translating the Gibbs sampling procedure above into MLM-based text sampling, we get:

MLM-Based Random Sampling
The initial sentence is $\boldsymbol{x}_0=(x_{0,1},x_{0,2},\cdots,x_{0,l})$, and the sentence at time $t$ is $\boldsymbol{x}_t=(x_{t,1},x_{t,2},\cdots,x_{t,l})$.
We sample a new sentence $\boldsymbol{x}_{t+1}$ via the following procedure:
1. Uniformly sample an index $i$ from $1,2,\cdots,l$, and replace the token at position $i$ with [MASK], obtaining the sequence $\boldsymbol{x}_{t,-i}=(x_{t,1},\dots,x_{t,i-1},\text{[MASK]},x_{t,i+1},\cdots,x_{t,l})$;
2. Feed $\boldsymbol{x}_{t,-i}$ into the MLM model and compute the probability distribution at position $i$, denoted $p_{t+1}$;
3. Sample a token from $p_{t+1}$, denoted $y$;
4. Replace the $i$-th token of $\boldsymbol{x}_{t}$ with $y$ to obtain $\boldsymbol{x}_{t+1}$.

Readers may have noticed that this sampling procedure can only sample sentences of a fixed length — it never changes the sentence length. That's indeed the case, because Gibbs sampling can only sample from within a single distribution, and sentences of different lengths in fact belong to different distributions, which theoretically have no overlap. It's just that when we usually build language models, we directly use an autoregressive model to uniformly model the distribution of sentences of all lengths, so we don't notice the fact that "sentences of different lengths actually belong to different probability distributions."

Of course, it's not impossible to work around this. The original paper BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model points out that we can set the initial sentence to be a sequence consisting entirely of [MASK] tokens. This way, we can first randomly sample a length $l$, and then start the Gibbs sampling process with $l$ [MASK] tokens as the initial sentence, thereby obtaining sentences of varying lengths.

Reference Code

With a ready-made MLM model in hand, implementing the Gibbs sampling procedure above is actually quite simple. Below is reference code implemented with bert4keras:

Gibbs sampling reference code: basic_gibbs_sampling_via_mlm.py

Here are some examples:

Initial sentence:
Science and technology are the primary productive force.
Sampling results:
How's the unboxing of the Honor laptop?
What to do if WeChat chat history is useless?
What to do if the browser can't be installed?
How to use the epf converter a7l?
What to do if the browser isn't installed?
How to use Honor laptop charging?
What to do if asp.net won't open?
What to do if the browser isn't installed?
What to do if the browser can't be restarted?
How to use the ro han ba conversion mac tv version?
Initial sentence:
Beijing reports 3 new local confirmed cases and 1 asymptomatic case.
Sampling results:
Macau recorded 233 cases of H1N1 infection and 13 cases of radioactive contamination.
The celebration was a grand event involving academy painting and steelworkers in the creation.
After the celebration, the Ghibli platform's other games also joined in.
Clinical trials found that g chromosomes mostly originate from gastrointestinal infections.
Clinical trials found that people typically genuinely enjoy clitoral pleasure.
The celebration mode is more in sync with the celebration on the Ghibli platform's other games.
The celebration mode has been updated and celebrated on Ghibli and other games.
Macau recorded 20 cases of H1N1 infection, 2 cases of radioactive contamination.
Clinical trials found that women's chromosomes commonly originate from gastrointestinal infections.
Clinical trials found that 90% of infection cases were type-m gastrointestinal infections.
Initial sentence:
9 consecutive [MASK]
Sampling results:
You act just like your mom every day!
That night, everything ahead was a hazy white.
Layer upon layer, lush and verdant green.
The kindergarten wants to start a business.
How exactly can one become an official?
Hello, teacher, hello, classmates!
Clouds upon clouds of mountains, both hazy and vast.
Plum rains, misty haze outside the window.
At that time, everything ahead was a hazy white.
The cake-cutting was really great!

The experiments here used Google's open-source Chinese BERT base model. As you can see, the sampled sentences are fairly diverse and have a certain degree of readability, which is already pretty good for a base-sized model.

For continuous [MASK] tokens as the initial input, repeated experiments can produce quite different results:

Initial sentence:
17 consecutive [MASK]
Sampling results:
What should other facial paralysis patients eat? What's good to eat for facial paralysis?
How should pediatric facial paralysis be treated? What medicine to take for facial paralysis?
How should facial paralysis in young children be treated? What's good to eat for facial paralysis?
What causes headaches in children? What disease is urticaria?
What should other facial paralysis patients eat, what's good to eat for other facial paralysis?
How do you actually fit the sanitary ware in, and how do you connect the faucet?
What's good to eat for other facial paralysis, what's good to eat for other facial paralysis?
What causes headaches in children? What disease is urticaria?
How do you actually fit the kitchen cabinet in, how do you insert the faucet?
Otherwise, if the kitchen cupboard can't be found, what do you do about the hot water faucet?
Initial sentence:
17 consecutive [MASK]
Sampling results:
Please check the tweet at the link below.
Please use the system we plan to operate in this area.
There are two crocar specialty stores in the area, please check the information!
Our site adopts a system that suits discounts!
The same work uses a superficial system.
The manufacturer's products genuinely use the system.
Please use the system available on the bulletin board.
The residence uses a production system.
Please check the address level list of Air Wear.
Please check the tweet's support at the link below.

Quite remarkable — Japanese text even gets sampled out, and having run it through Baidu Translate, the author found this Japanese text to be reasonably readable. On one hand, this illustrates the diversity of random sampling results; on the other, it also shows that Google's Chinese BERT wasn't thoroughly denoised, and its training corpus must have contained a fair amount of non-Chinese, non-English text mixed in.

Some Food for Thought

A while back, Google, Stanford, OpenAI and others jointly published a paper, Extracting Training Data from Large Language Models, which showed that language models like GPT2 are fully capable of reproducing (leaking) their training data. This isn't hard to understand, since a language model is essentially reciting sentences at heart. And MLM-based Gibbs sampling shows that this issue isn't unique to explicit language models like GPT2 — bidirectional language models like MLMs suffer from it too. We can already see hints of this in the sampling examples above: for instance, the fact that Japanese text got sampled suggests the original corpus wasn't especially well denoised, and starting from "Beijing reports 3 new local confirmed cases and 1 asymptomatic case," we sampled some H1N1-related results, which reflects the era in which the training corpus was collected. All of this implies that if you don't want your open-sourced model to leak your private data, you need to do a thorough job of cleaning the pretraining corpus.

There's also another bit of gossip worth mentioning about the paper BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model: the original paper claimed that the MLM model is a Markov Random Field, but this turns out not to be true. The authors later clarified this on their own homepage — interested readers can check out BERT has a Mouth and must Speak, but it is not an MRF. In short: using MLM for random sampling works fine, but it doesn't quite correspond to a Markov Random Field.

Summary

This post introduced random text sampling based on BERT's MLM, which is essentially a natural application of Gibbs sampling. Overall, this is a fairly simple example. For readers already familiar with Gibbs sampling, there's almost no technical difficulty here; and for readers who aren't yet very familiar with Gibbs sampling, this concrete example provides a good opportunity to further understand how the Gibbs sampling procedure works.

English translation of a post from 科学空间 | Scientific Spaces by 苏剑林. Original: https://kexue.fm/archives/8119
Translated automatically with claude-sonnet-5; all equations are reproduced verbatim from the source. Copyright remains with the original author.