Dataset

A dataset is a systematically organized collection of data, typically structured in a way that facilitates analysis and interpretation. In its most common form, a dataset can be represented as a table, where each row corresponds to an individual observation (for example, a person, a product, or a measurement), and each column corresponds to a specific attribute or variable (such as age, income, temperature, or category).

For instance, in a dataset describing real estate properties, each row would represent a unique house, while columns would include characteristics such as price, surface area, number of bedrooms, geographic location, and year of construction. This structure allows for efficient data manipulation, comparison, and statistical analysis.

Datasets are not limited to tabular structures, they can be structured, semi-structured, or unstructured, depending on how the data are formatted and stored. Structured datasets, such as those found in relational databases, follow a well-defined schema with clearly identified variables. Semi-structured datasets, like JSON or XML files, contain data that follow a flexible but partially organized format. Unstructured datasets, such as images, audio recordings, or free-form text, lack a predefined organizational model but can still be analyzed using advanced computational methods, including machine learning and natural language processing.

In contemporary data science, the quality and representativeness of a dataset are critical. Biases, missing values, or sampling errors can significantly influence the results of any analysis or model built upon it. Thus, the process of data collection, cleaning, and validation plays a fundamental role in ensuring the reliability of subsequent findings.


Distribution

A distribution describes how the values of a variable within a dataset are dispersed or arranged. It represents the frequency or probability with which each possible value occurs, providing insight into the underlying structure of the data.

Several types of distributions are commonly encountered in data analysis. The normal distribution, often represented by the classic bell-shaped curve, is characterized by symmetry around the mean, where most observations cluster near the center and probabilities decrease toward the tails. Many natural and social phenomena—such as human height or measurement errors—tend to approximate this distribution.

In contrast, a uniform distribution assumes that all outcomes are equally likely, as in the case of a fair die roll. Other distributions exhibit asymmetry, such as skewed distributions, where data are concentrated on one side, producing a long tail on the other. Income distributions, for example, are typically right-skewed, with a large concentration of moderate values and a few extremely high ones.

Understanding the shape and parameters of a distribution is essential for selecting the most suitable statistical techniques and for modeling uncertainty within data-driven systems.


The Connection Between Dataset and Distribution

While a dataset represents the actual, finite collection of observations you have gathered, the distribution reflects the theoretical or empirical pattern that these observations follow. In statistical terms, a dataset can be considered a sample drawn from an underlying population distribution. By studying the dataset’s empirical distribution, analysts aim to infer characteristics of this broader, often unobservable, population.

This connection is central to all inferential and predictive analyses. Understanding the distribution underlying a dataset allows researchers to make more accurate predictions, design more robust models, and detect deviations or anomalies that may indicate errors, biases, or emerging trends. Ultimately, the interplay between dataset and distribution forms the foundation of both statistical reasoning and data-driven decision-making.

In the following, a Database Management System (DBMS) is employed to create a small dataset for analysis. The main goal is to compute the univariate distribution of up to three selected variables (k ≤ 3). Univariate distribution allows us to describe how the values of a single variable are distributed, showing the frequency of each value or range.

Create a table called Students, which will include the following fields: id, name, age, course, examCount and averageGrade.

Insert the values

The next step is to compute the k univariate distributions, that describes how the values of a single variable are distributed, showing the frequency of each value or range. The following section present each distribution along with its corresponding results.

The first one:

The second one:

The third one:

In the last one, there is a bivariate distribution shows how two variables are related and how their values occur together. It helps identify patterns, correlations, or associations between the variables.

2. Caesar Cipher

This phase involves computing the distribution of each individual letter of the English alphabet within the text. The analyzed text is shown after the code used for computation.

function distribEachLetter(text) {
  const lowerCase = text.toLowerCase().replace(/[^a-z]/g, '');
  const distribution = {};

  for (const char of lowerCase) {
    distribution[char] = (distribution[char] || 0) + 1;
  }

  const total = lowerCase.length;
  for (const char in distribution) {
    distribution[char] = {
      count: distribution[char],
      percentage: ((distribution[char] / total) * 100).toFixed(2) + '%'
    };
  }
  return distribution;
}

The following graph displays the letter distribution:

The code presented below encrypts the text using the Caesar cipher with a shift value of 8. The Caesar cipher, one of the simplest and oldest encryption techniques, is a substitution cipher in which each letter in the plaintext is systematically replaced by a letter located a fixed number of positions further in the alphabet. With a shift of 8, each character is displaced 8 positions forward (for instance, ‘A’ is replaced by ‘I’, ‘B’ by ‘J’, and so forth). Upon reaching the end of the alphabet, the substitution wraps around to the beginning.

function caesarCipher(text, shift = 8) {
  const alphabet = 'abcdefghijklmnopqrstuvwxyz';
  const upperAlphabet = alphabet.toUpperCase();
  let result = '';

  for (let char of text) {
    if (alphabet.includes(char)) {
      const newIndex = (alphabet.indexOf(char) + shift) % 26;
      result += alphabet[newIndex];
    } else if (upperAlphabet.includes(char)) {
      const newIndex = (upperAlphabet.indexOf(char) + shift) % 26;
      result += upperAlphabet[newIndex];
    } else {
      result += char;
    }
  }

  return result;
}
The Ciphertext can be found here

Gwc eqtt zmrwqkm bw pmiz bpib vw lqaiabmz pia ikkwuxivqml bpm kwuumvkmumvb wn iv mvbmzxzqam epqkp gwc pidm zmoizlml eqbp ackp mdqt nwzmjwlqvoa. Q izzqdml pmzm gmabmzlig, ivl ug nqzab bias qa bw iaaczm ug lmiz aqabmz wn ug emtnizm ivl qvkzmiaqvo kwvnqlmvkm qv bpm ackkmaa wn ug cvlmzbisqvo. Q iu itzmilg niz vwzbp wn Twvlwv, ivl ia Q eits qv bpm abzmmba wn Xmbmzajczop, Q nmmt i kwtl vwzbpmzv jzmmhm xtig cxwv ug kpmmsa, epqkp jzikma ug vmzdma ivl nqtta um eqbp lmtqopb. Lw gwc cvlmzabivl bpqa nmmtqvo? Bpqa jzmmhm, epqkp pia bzidmttml nzwu bpm zmoqwva bweizla epqkp Q iu ildivkqvo, oqdma um i nwzmbiabm wn bpwam qkg ktquma. Qvaxqzqbml jg bpqa eqvl wn xzwuqam, ug liglzmiua jmkwum uwzm nmzdmvb ivl dqdql. Q bzg qv diqv bw jm xmzacilml bpib bpm xwtm qa bpm amib wn nzwab ivl lmawtibqwv; qb mdmz xzmamvba qbamtn bw ug quioqvibqwv ia bpm zmoqwv wn jmicbg ivl lmtqopb. Bpmzm, Uizoizmb, bpm acv qa nwz mdmz dqaqjtm, qba jzwil lqas rcab asqzbqvo bpm pwzqhwv ivl lqnncaqvo i xmzxmbcit axtmvlwcz. Bpmzm—nwz eqbp gwcz tmidm, ug aqabmz, Q eqtt xcb awum bzcab qv xzmkmlqvo vidqoibwza—bpmzm avwe ivl nzwab izm jivqapml; ivl, aiqtqvo wdmz i kitu ami, em uig jm einbml bw i tivl aczxiaaqvo qv ewvlmza ivl qv jmicbg mdmzg zmoqwv pqbpmzbw lqakwdmzml wv bpm pijqbijtm otwjm. Qba xzwlckbqwva ivl nmibczma uig jm eqbpwcb mfiuxtm, ia bpm xpmvwumvi wn bpm pmidmvtg jwlqma cvlwcjbmltg izm qv bpwam cvlqakwdmzml awtqbclma. Epib uig vwb jm mfxmkbml qv i kwcvbzg wn mbmzvit tqopb? Q uig bpmzm lqakwdmz bpm ewvlzwca xwemz epqkp ibbzikba bpm vmmltm ivl uig zmoctibm i bpwcaivl kmtmabqit wjamzdibqwva bpib zmycqzm wvtg bpqa dwgiom bw zmvlmz bpmqz ammuqvo mkkmvbzqkqbqma kwvaqabmvb nwz mdmz. Q apitt aibqibm ug izlmvb kczqwaqbg eqbp bpm aqopb wn i xizb wn bpm ewztl vmdmz jmnwzm dqaqbml, ivl uig bzmil i tivl vmdmz jmnwzm quxzqvbml jg bpm nwwb wn uiv. Bpmam izm ug mvbqkmumvba, ivl bpmg izm acnnqkqmvb bw kwvycmz itt nmiz wn livomz wz lmibp ivl bw qvlckm um bw kwuumvkm bpqa tijwzqwca dwgiom eqbp bpm rwg i kpqtl nmmta epmv pm mujizsa qv i tqbbtm jwib, eqbp pqa pwtqlig uibma, wv iv mfxmlqbqwv wn lqakwdmzg cx pqa vibqdm zqdmz. Jcb acxxwaqvo itt bpmam kwvrmkbczma bw jm nitam, gwc kivvwb kwvbmab bpm qvmabquijtm jmvmnqb epqkp Q apitt kwvnmz wv itt uivsqvl, bw bpm tiab omvmzibqwv, jg lqakwdmzqvo i xiaaiom vmiz bpm xwtm bw bpwam kwcvbzqma, bw zmikp epqkp ib xzmamvb aw uivg uwvbpa izm zmycqaqbm; wz jg iakmzbiqvqvo bpm amkzmb wn bpm uiovmb, epqkp, qn ib itt xwaaqjtm, kiv wvtg jm mnnmkbml jg iv cvlmzbisqvo ackp ia uqvm. Bpmam zmntmkbqwva pidm lqaxmttml bpm ioqbibqwv eqbp epqkp Q jmoiv ug tmbbmz, ivl Q nmmt ug pmizb otwe eqbp iv mvbpcaqiau epqkp mtmdibma um bw pmidmv, nwz vwbpqvo kwvbzqjcbma aw uckp bw bzivycqttqam bpm uqvl ia i abmilg xczxwam—i xwqvb wv epqkp bpm awct uig nqf qba qvbmttmkbcit mgm. Bpqa mfxmlqbqwv pia jmmv bpm nidwczqbm lzmiu wn ug miztg gmiza. Q pidm zmil eqbp izlwcz bpm ikkwcvba wn bpm dizqwca dwgioma epqkp pidm jmmv uilm qv bpm xzwaxmkb wn izzqdqvo ib bpm Vwzbp Xikqnqk Wkmiv bpzwcop bpm amia epqkp aczzwcvl bpm xwtm. Gwc uig zmumujmz bpib i pqabwzg wn itt bpm dwgioma uilm nwz xczxwama wn lqakwdmzg kwuxwaml bpm epwtm wn wcz owwl Cvktm Bpwuia’ tqjzizg. Ug mlckibqwv eia vmotmkbml, gmb Q eia xiaaqwvibmtg nwvl wn zmilqvo. Bpmam dwtcuma emzm ug abclg lig ivl vqopb, ivl ug niuqtqizqbg eqbp bpmu qvkzmiaml bpib zmozmb epqkp Q pil nmtb, ia i kpqtl, wv tmizvqvo bpib ug nibpmz’a lgqvo qvrcvkbqwv pil nwzjqllmv ug cvktm bw ittwe um bw mujizs qv i aminizqvo tqnm. Bpmam dqaqwva nilml epmv Q xmzcaml, nwz bpm nqzab bqum, bpwam xwmba epwam mnncaqwva mvbzivkml ug awct ivl tqnbml qb bw pmidmv. Q itaw jmkium i xwmb ivl nwz wvm gmiz tqdml qv i xizilqam wn ug wev kzmibqwv; Q quioqvml bpib Q itaw uqopb wjbiqv i vqkpm qv bpm bmuxtm epmzm bpm viuma wn Pwumz ivl Apismaxmizm izm kwvamkzibml. Gwc izm emtt ikyciqvbml eqbp ug niqtczm ivl pwe pmidqtg Q jwzm bpm lqaixxwqvbumvb. Jcb rcab ib bpib bqum Q qvpmzqbml bpm nwzbcvm wn ug kwcaqv, ivl ug bpwcopba emzm bczvml qvbw bpm kpivvmt wn bpmqz miztqmz jmvb. Aqf gmiza pidm xiaaml aqvkm Q zmawtdml wv ug xzmamvb cvlmzbisqvo. Q kiv, mdmv vwe, zmumujmz bpm pwcz nzwu epqkp Q lmlqkibml ugamtn bw bpqa ozmib mvbmzxzqam. Q kwuumvkml jg qvczqvo ug jwlg bw pizlapqx. Q ikkwuxivqml bpm epitm-nqapmza wv amdmzit mfxmlqbqwva bw bpm Vwzbp Ami; Q dwtcvbizqtg mvlczml kwtl, niuqvm, bpqzab, ivl eivb wn atmmx; Q wnbmv ewzsml pizlmz bpiv bpm kwuuwv aiqtwza lczqvo bpm lig ivl lmdwbml ug vqopba bw bpm abclg wn uibpmuibqka, bpm bpmwzg wn umlqkqvm, ivl bpwam jzivkpma wn xpgaqkit akqmvkm nzwu epqkp i vidit ildmvbczmz uqopb lmzqdm bpm ozmibmab xzikbqkit ildivbiom. Beqkm Q ikbcittg pqzml ugamtn ia iv cvlmz-uibm qv i Ozmmvtivl epitmz, ivl ikycqbbml ugamtn bw iluqzibqwv. Q ucab wev Q nmtb i tqbbtm xzwcl epmv ug kixbiqv wnnmzml um bpm amkwvl lqovqbg qv bpm dmaamt ivl mvbzmibml um bw zmuiqv eqbp bpm ozmibmab mizvmabvmaa, aw ditcijtm lql pm kwvaqlmz ug amzdqkma. Ivl vwe, lmiz Uizoizmb, lw Q vwb lmamzdm bw ikkwuxtqap awum ozmib xczxwam? Ug tqnm uqopb pidm jmmv xiaaml qv miam ivl tcfczg, jcb Q xzmnmzzml otwzg bw mdmzg mvbqkmumvb bpib emitbp xtikml qv ug xibp. Wp, bpib awum mvkwczioqvo dwqkm ewctl ivaemz qv bpm innqzuibqdm! Ug kwcziom ivl ug zmawtcbqwv qa nqzu; jcb ug pwxma ntckbcibm, ivl ug axqzqba izm wnbmv lmxzmaaml. Q iu ijwcb bw xzwkmml wv i twvo ivl lqnnqkctb dwgiom, bpm mumzomvkqma wn epqkp eqtt lmuivl itt ug nwzbqbclm: Q iu zmycqzml vwb wvtg bw ziqam bpm axqzqba wn wbpmza, jcb awumbquma bw acabiqv ug wev, epmv bpmqza izm niqtqvo. Bpqa qa bpm uwab nidwczijtm xmzqwl nwz bzidmttqvo qv Zcaaqi. Bpmg ntg ycqkstg wdmz bpm avwe qv bpmqz atmloma; bpm uwbqwv qa xtmiaivb, ivl, qv ug wxqvqwv, niz uwzm iozmmijtm bpiv bpib wn iv Mvotqap abiomkwikp. Bpm kwtl qa vwb mfkmaaqdm, qn gwc izm ezixxml qv ncza—i lzmaa epqkp Q pidm itzmilg ilwxbml, nwz bpmzm qa i ozmib lqnnmzmvkm jmbemmv eitsqvo bpm lmks ivl zmuiqvqvo amibml uwbqwvtmaa nwz pwcza, epmv vw mfmzkqam xzmdmvba bpm jtwwl nzwu ikbcittg nzmmhqvo qv gwcz dmqva. Q pidm vw iujqbqwv bw twam ug tqnm wv bpm xwab-zwil jmbemmv Ab. Xmbmzajczop ivl Izkpivomt. Q apitt lmxizb nwz bpm tibbmz bwev qv i nwzbvqopb wz bpzmm emmsa; ivl ug qvbmvbqwv qa bw pqzm i apqx bpmzm, epqkp kiv miaqtg jm lwvm jg xigqvo bpm qvaczivkm nwz bpm wevmz, ivl bw mvoiom ia uivg aiqtwza ia Q bpqvs vmkmaaizg iuwvo bpwam epw izm ikkcabwuml bw bpm epitm-nqapqvo. Q lw vwb qvbmvl bw aiqt cvbqt bpm uwvbp wn Rcvm; ivl epmv apitt Q zmbczv? Ip, lmiz aqabmz, pwe kiv Q ivaemz bpqa ycmabqwv? Qn Q ackkmml, uivg, uivg uwvbpa, xmzpixa gmiza, eqtt xiaa jmnwzm gwc ivl Q uig ummb. Qn Q niqt, gwc eqtt amm um ioiqv awwv, wz vmdmz. Nizmemtt, ug lmiz, mfkmttmvb Uizoizmb. Pmidmv apwemz lwev jtmaaqvoa wv gwc, ivl aidm um, bpib Q uig ioiqv ivl ioiqv bmabqng ug ozibqbclm nwz itt gwcz twdm ivl sqvlvmaa. Gwcz innmkbqwvibm jzwbpmz, Z. Eitbwv

The following graph displays the letter distribution of the ciphertext above:

The final graph

The final graph displays both the letter distribution in the plaintext and the ciphertext, allowing for a direct comparison between the two. By identifying the highest frequency values to minimize errors, it can be observed that ‘E’ is the most frequently repeated letter in the plaintext, while ‘M’ is the most frequent in the ciphertext. Based on this observation, the shift can be determined to be 8 positions, as ‘M’ is located 8 positions after ‘E’ in the alphabet.

Decoding without knowing the shift (Brute Force)

In order to decrypt the ciphertext without knowing the shift value in advance, thereby recovering the original plaintext and determining the shift that was applied, the following function may be employed.

function bruteForce(ciphertext) {
  for (let i = 0; i < 26; i++) {
    const decoded = caesarCipher(ciphertext, (26 - i) % 26);
    console.log(`shift ${i}: ${decoded}`);
  }
}

Decoding using “Language Distribution“

To decrypt the ciphertext without prior knowledge of the shift value, frequency analysis is employed as the primary cryptanalytic technique. This method relies on comparing the letter frequency distribution of the encrypted text with the known statistical distribution of letters in standard English language texts.

Standard English Letter Frequencies (Top 10):

LetterFrequency
E12.702%
T9.056%
A8.167%
O7.507%
I6.966%
N6.749%
S6.327%
H6.094%
R5.987%
D4.253%

Ciphertext Letter Frequencies (Top 10):

LetterCountPercentage
M72413.33%
B4618.49%
I4528.32%
V3857.09%
Q4207.73%
W3706.81%
A3356.17%
Z3396.24%
P2895.32%
L2314.25%

Comparative Analysis of the Five Most Frequent Letters:

A systematic comparison of the five most frequent letters in each distribution reveals a consistent pattern:

  • E (12.70%) → M (13.33%): The alphabetic distance from E to M is 8 positions
  • T (9.06%) → B (8.49%): T shifted by 8 positions (modulo 26) yields B
  • A (8.17%) → I (8.32%): The mapping from A to I corresponds to a displacement of 8 positions
  • O (7.51%) → W (6.81%): O advanced by 8 positions results in W
  • I (6.97%) → Q (7.73%): I shifted forward by 8 positions produces Q

The analysis demonstrates a uniform shift of 8 positions across all high-frequency letters. This consistency across multiple data points provides strong evidence that the Caesar cipher implementation employed a shift value of 8. Consequently, the original plaintext can be recovered by applying the inverse transformation, shifting each letter backward by 8 positions in the alphabet.