Microsoft Forms Did Me Dirty
Validating Third-Party Tool Data Analysis
Summary: After MS Forms produced a questionable winner from a 32-person force-ranking exercise, I exported the raw data and compared the default output with Borda scoring, pairwise wins, and the Schulze method. The analysis showed that different methods can produce different winners, which is why UX researchers need enough statistical knowledge to recognize when a tool’s default calculation does not match the question they are trying to answer.
A couple of weeks ago, I used MS Forms to run a stack-ranking exercise at work. 32 participants ranked 8 items from most important to least important. When I opened the results, Forms had already calculated the rankings, as it should, and identified a winner.
The result did not look wrong to the untrained eye, but within minutes of looking at the raw data, I knew something was off. For example, the item Forms ranked 1st appeared consistently near the top of everyone’s lists, but another item received far more first-place rankings and performed better when compared directly with most of the other options. So I knew what I needed to do. I needed to determine the most appropriate way to reanalyze the data because Forms’ default analysis sucked for this scenario. Thanks Microsoft.
Study Design Constraints
Let’s back up a little. Stack ranking was the best approach here because the stakeholders needed to understand relative preference across a short list of comparable items. Think of a study asking users to prioritize 8 common tasks they might complete through a self-service portal:
Check the status of a request
Update account information
Download a document
Make a payment
Contact support
Review recent activity
Manage notification settings
Add or remove an authorized user
These are concise, distinct tasks that participants can reasonably compare. Ranking them forces people to make tradeoffs that a standard rating scale might not reveal. Someone could reasonably rate all 8 tasks as important, but a forced ranking shows which ones they would prioritize when everything cannot be treated equally.
Another factor was the deadline was super tight, aren’t they all. Hahahaha. Based on our previous recruiting and survey conversion rates, I expected between 30 and 50 completed responses within the timeline we had to work in. I would have preferred a larger sample because it would have provided greater precision and made smaller differences between items easier to interpret. A larger sample would also have reduced the likelihood that a few individual responses could change the final order.
The 32 responses were still very helpful, relevant and statistically valid within the guardrails established for the study. I presented the findings as directional evidence, avoided claims about small differences and did not use the results for subgroup comparisons. And here’s the important takeaway, I explained those limits before launching the survey. The stakeholders agreed upfront that the study would still provide enough evidence to guide the immediate decision.
This upfront conversation is a very mature research step that many UX Researchers flat out skip. Getting agreement on the strength of the evidence before collecting it reduces arguments about the method during the final readout. It also gives the research team credibility because the standard for interpreting the results is established before anyone knows which item will win.
The Real World Requires Judgment
MS Forms produced an aggregate ranking automatically. Based on the output, it appeared to be applying a basic positional weighting method. Items received more value when participants placed them near the top and progressively less value in each lower position.
I could not confirm the exact formula because Forms did not explain the calculation and I don’t have the patience to go look it up. The charts Forms pooped out weren’t necessarily wrong but they didn't pass the sniff test. This type of scoring can work well when you want to find the item that most people ranked near the top. But that does not always mean it was the item people preferred most strongly or chose over the other options most often. This kind of nuanced and expert judgement is something I see missed by most UXers in the field today.
UX researchers do not need to be as statistically fluent as data analysts. But, we do need enough basic statistical understanding to recognize when a default analysis may not fit the question, the sample or the decision being made.
Third-party research platforms have to select a default calculation that can work across many use cases. Even a product from a company such as Microsoft cannot determine the most appropriate interpretation for every study.
This type of discernment will become more important as automated and AI-generated analysis becomes common. The output may be polished and mathematically correct while still answering a different question from the one the research was designed to address.
Testing Alternatives
So what did I do about it? First, I downloaded the individual responses into Excel and kept one row for each participant. I converted each answer into a numerical ranking from one through eight and confirmed that every respondent had used each rank once. I then tested two alternative approaches.
⚠️ Disclaimer: I did not immediately know how to do all of this by accident. I spent almost 9 years working for Minitab, one of the most sophisticated desktop statistical analysis products available. That experience helped me recognize that the default calculation might not fit this particular dataset and research question.
For those who are not confident they would have spotted the issue, there is no need to worry. I tested the same question with 3 different large language models, and all three gave me useful responses. This is exactly the kind of task LLMs can be great at: helping you identify possible analysis methods, understand what each one measures and determine when a tool’s default output deserves a closer look.
Here is the prompt I used:
I conducted a forced-ranking survey in which 32 participants ranked 8 items from most important to least important. The survey platform automatically produced an overall ranking, but it does not explain the exact calculation it used.
After reviewing the individual responses, I noticed that the item ranked first by the platform was consistently placed near the top, while another item received substantially more first-place rankings and appeared to beat most of the other items in direct comparisons.
Help me determine the most appropriate way to analyze this dataset. Explain, in straightforward language.
The first method I tried was a combined Borda scoring with pairwise wins. Borda assigns points according to where an item appears in each participant’s ranking. Pairwise analysis compares every item directly with each of the other seven items.
For example, the analysis would compare “check the status of a request” with “contact support” and count how many participants placed each task higher. The same comparison would then be completed for every possible pair.
I used the number of pairwise wins as the primary measure and the Borda score as the tie-breaker. This approach emphasized direct preference while still recognizing items that participants placed consistently near the top.
The second approach was the Schulze method. Schulze also begins with pairwise comparisons, but it considers the strength of the paths connecting the items. This allows it to resolve situations where preferences form cycles and no item defeats every alternative cleanly.
Schulze is a strong method when the analysis needs to produce one formal winner from complicated pairwise results. Its practical weakness is explainability. The calculation is more difficult for stakeholders to review, understand and reproduce.
Different Methods, Different Stories
The following example uses fictional tasks and a reconstructed response pattern. It illustrates the analytical issue without using any company information or reproducing the actual study content.
Something like this:
Method #1: MS Forms
The default weighted method, the MS Form’s method, selected “check request status” because participants consistently placed it near the top. It had broad support and relatively few low rankings.
Method #2: Schulze
Schulze selected “contact support.” It received the most 1st votes and performed strongly in the direct comparisons. The method looked at how often participants ranked “contact support” above each of the other 7 tasks, then considered the strength of those head-to-head wins across the full set of comparisons. That gave it the strongest overall path through the rankings, even though it did not have the highest positional score.
Method #3: Borda and Pairwise Combo
The pairwise-plus-Borda approach produced a tie between “download a document” and “contact support.” Each defeated six of the other seven tasks. “Download a document” then won the Borda tie-break because participants placed it more consistently across the complete rankings.
A similar problem appears in the ForceRank.it article “Counting Votes Is Hard.” Four people ranked 9 possible topics for a technical presentation. Three people placed Item 1 first, but ForceRank’s original point-based scoring method selected Item 2. Item 2 performed consistently well across all four rankings, while Item 1 was the majority favorite but was placed last by one participant. ForceRank eventually replaced its original method with the Schulze method because the original calculation failed the Majority Criterion. Schulze evaluates the items through pairwise comparisons and produces a winner that better reflects the group’s overall preferences.
So in short, the original calculation had not failed. It just had used a different definition of winning in the same way MS Forms did.
Conclusion
For this study, I chose pairwise wins with Borda as the tie-breaker. That combo method was understandable, reproducible and appropriate for a directional sample of 32 respondents. It also made it possible to explain why the methods produced different results without presenting one calculation as universally correct.
I compared every item directly with the other seven and counted how many comparisons it won. When two items had the same number of wins, I used the Borda score as the tie-breaker. I also reported first-place votes and the full rank distributions so stakeholders could see the preference pattern behind the final order.
⚠️ Disclaimer: I would still use Schulze when the pairwise relationships contained complex cycles and the study required one formal winner. It solves that specific mathematical problem well.
MS Forms collected the responses and provided a reasonable default summary. My responsibility was to determine whether that summary matched the stakeholder’s question.
The most important analytical skill in this project was knowing when the default answer needed a second look. The tools will keep getting better, but our responsibility stays the same. We need to understand the method well enough to know when the default output does not fit the research question. I hope sharing this story was helpful. Thanks for reading.





I can't stand Microsoft products now a days... I wouldn't trust them to calculate anything past some spreadsheet totals lol!