How Dating Apps Use Behavior Data to Recommend Potential Matches

 

The least useful thing a dating app knows about you is what you said you wanted. The filters you set and the paragraph you wrote about your ideal partner are close to decorative. What drives the recommendations is the record of what you did, and those two things routinely disagree.

People report their preferences accurately and then behave differently, and psychologists have measured the size of that discrepancy under controlled conditions.

The Two Kinds of Preference

A 2008 speed-dating study asked participants to rate the importance of qualities in an ideal partner before the event. The stated answers produced the differences between men and women that everyone expects, with men rating physical attractiveness higher and women rating earning prospects higher.

Then the participants met each other. The preferences they had stated beforehand showed no correlation with who they actually wanted afterward. Later work in the same line found something sharper. Matching someone against their stated ideals does predict how they will rate a photograph, and the effect weakens to almost nothing once two people have spoken face to face.

That result explains the entire design of a recommendation system built on behavior. In the part of the process an app controls, which is a person looking at pictures, stated preferences retain some predictive value. The system is operating in exactly the regime where the weaker signal still works.

What Actually Gets Logged

In 2017 a French journalist used European data protection law to request everything one large platform held on her. The reply was 800 pages long.

It contained her Facebook likes, photographs pulled from an Instagram account she had already deleted, her education, the age-rank of the men she had shown interest in, the number of times she had connected, and the time and place of every conversation she had held on the service. It included the full text of 1,700 messages she had sent since 2013. Obtaining it required two formal complaints, dozens of emails, and months of waiting, with help from a data protection activist and a human rights lawyer.

Almost all of that file is a behavioral record.

Declared Intent at Signup

The distance between stated and revealed preference is widest where a user declares almost nothing when joining. A general platform receives an age range, a location radius, and a paragraph of free text, then spends the following months inferring everything of consequence from swipe records.

Services organized around one declared purpose begin from a different position. A platform for a single religious community, a service for people who want children inside two years, or a sugar baby website collects the load-bearing preference as a condition of entry, which leaves less of the profile to be reconstructed from behavior after the fact.

Implicit Signals

Swipes are the obvious input and the least interesting one. The finer signals come from timing and attention.

How long a profile stays on screen before a decision separates a considered yes from a reflexive one. How quickly a person replies to a first message, and how long that reply is, predicts how invested they are in a given match. Which photograph in a set holds attention longest indicates something the user would struggle to articulate. Session times matter too, since a person who opens the app at 11pm on weeknights is being modeled differently from one who opens it on Sunday mornings.

None of these require a user to answer anything. They accumulate on their own, with no participation from the person producing them, which is the property that makes them valuable to a recommendation system and uncomfortable to think about. Services are obliged to disclose the behavioral data they gather and pass on, and reading that disclosure is the fastest way to see how much privacy at risk actually means in practice.

Third-Party Data Sharing

The collection is only half of it. In 2020 the Norwegian Consumer Council examined 10 popular apps and reported that user data was being transmitted to at least 135 separate third-party companies involved in advertising or behavioral profiling. Data-sharing across the apps it studied had run out of control, and none of the apps or the firms receiving the data met the legal conditions for valid consent.

One service in that report was found to be passing GPS coordinates, IP addresses, ages, and genders to outside firms for ad targeting. The Norwegian data protection authority fined it 65 million kroner, roughly €6.3 million, for sharing user data without consent. A separate regulator opened an investigation into a larger platform’s handling of user data the same year.

The behavioral profile built to recommend matches has a second life as an advertising asset, and the second use is the one that funds the first.

Training on Its Own Output

A model trained on behavior has a structural weakness. It learns from choices that the model itself produced.

If the system shows a user a narrow band of profiles, the user can only choose from that band, and those choices confirm the model’s estimate. There is no correction available from inside the loop, because the app never observes what would have happened had it shown something else. The mechanism is the same one that produces filter bubbles in a news feed, where a system optimizing for relevance keeps narrowing the range of what a person is shown. This is why users describe their recommendations as becoming repetitive after a few weeks of heavy use, and why deleting an account and starting over sometimes produces visibly different results from adjusting the filters on the existing one.

The Photograph Boundary

Everything described here operates on a person evaluating images and text on a screen. The 2008 research marks the boundary precisely. Preference-matching predicts photograph ratings and loses nearly all of its force once two people meet.

A recommendation engine is therefore optimized for the wrong end of the process. It is very good at predicting which profile a user will tap and has no measurable ability to predict which dinner will go well. Users who treat a high volume of matches as evidence of good matching are reading a metric the system is good at and mistaking it for one it cannot measure.

The Numbers Worth Keeping

One journalist’s file came to 800 pages, covering 1,700 messages and the location of every conversation she had held since 2013. Ten apps in one audit were feeding data to 135 third parties. A user deciding how much to invest in a recommendation should hold those figures alongside the one number the research supports, which is zero correlation between the preferences people state and the people they actually want after meeting them.