Malmö University Publications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
En jämförelse av maskininlärningsalgoritmer för uppskattning av cykelflöden baserat på cykelbarometer- och väderdata
Malmö högskola, Faculty of Technology and Society (TS).
Malmö högskola, Faculty of Technology and Society (TS).
2016 (Swedish)Independent thesis Basic level (degree of Bachelor), 10 credits / 15 HE creditsStudent thesis
Abstract [sv]

Kontext. Maskininlärningsalgoritmer kan användas för att göra förutsägelser baserat på en mängd data. Vi använder oss utav data ifrån en cykelbarometer lokaliserad vid en cy- kelväg i Malmö i vår forskning. Denna barometer räknar antalet förbipasserande cyklar per dag. Tillsammans med väderdata, som består av temperatur och nederbörd, jämför vi precisionen hos algoritmer för uppskattning av antalet cyklister. I denna studie imple- menterar vi och testar en mängd olika maskininlärningsalgoritmer som finns tillgängliga i programvaran Weka. Vi tar hjälp av tidigare forskning inom ämnet för att identifiera vilka algoritmer som lämpar sig bäst för vår typ av data. Vi väljer sedan ut de tre algoritmer med bäst träffsäkerhet och undersöker dessa närmare. Mål. Målet med studien är att vi ska få fram vilken maskininlärningsalgoritm som ger det mest tillförlitliga resultatet för att uppskatta antalet cyklister med hjälp av vår cykel- barometer- och väderdata. Metoder. Vi bearbetar datan ifrån cykelbarometern och väderstationen för att filtrera bort dagar som kan förvränga resultatet. Exempel på data som vi filtrerar bort är helgdagar och skollov. Med den filtrerade datan implementerar vi ett flertal maskininlärningsalgorit- mer för att uppskatta antalet cyklister som kommer att passera barometern under en nära framtid. Resultaten ifrån algoritmerna använder vi för att jämföra och se vilken algoritm som ger den mest tillförlitliga uppskattningen för den aktuella tillämpningen. Resultat. Enligt våra resultat är Random SubSpace och Bagging de överlägsna algorit- merna för att uppskatta cykelflöde. I samtliga av våra experiment åstadkommer dessa två bättre resultat än övriga algoritmer som finns tillgängliga i Weka. Resultaten därefter skil- jer sig från experiment till experiment men i genomsnitt är Wekas REPTree-algoritm den tredje mest precisa. Variabeln som bidrar mest till vår uppskattning av antalet cyklister är datum. Utan denna variabel reduceras korrelationen till hälften för samtliga algoritmer. När vi avlägsnar temperatur-variabeln presterar däremot algoritmerna bättre genom att ge högre korrelation. Analys. Vi har hittat en korrelation mellan datum och cykelflöden samt kunnat förutsäga cykelflöden beroende på datum och väder. Vi förväntade oss inte att variabeln temperatur gör det svårare för algoritmer att uppskatta antal cyklister. Vi antar att detta beror på att människor väljer att cykla efter datum istället för temperatur.

Abstract [en]

Context. Machine Learning Algorithms can be used to make predictions based on a va- riety of data. We use data from a bicycle barometer located at a bike path in Malmö in our research. This barometer counts the number of passing bikes per day. Together with weather data, consisting of temperature and precipitation, we compare the accuracy of the algorithms to estimate the number of cyclists. In this study we implement and test a variety of machine learning algorithms that are available in the software Weka. We rely on previous research in order to identify which algorithms are best suited for our type of data. We will then select the three algorithms with the best accuracy and examine them closer. Goal. The goal of the study is to identify the machine learning algorithm that provides the most reliable results to estimate the number of cyclists using our bicycle barometer- and weather data. Methods. We process the data from the bicycle barometer and weather station to filter out days that can distort the results. Examples of data that we filter out are public holidays and school holidays. With the filtered data we implement three different machine learning algorithms to estimate the number of bicyclists who will pass the barometer in the near future. The results from the algorithms are then used to compare and see which algorithm that makes the most reliable estimate of the current application. Results. According to our results, the Random SubSpace and Bagging methods are the superior algorithms to estimate the cycle flow. These algorithms provide the best results in all of our experiments. The results differ beyond those two algorithms but on average Wekas REPTree algorithm is the third most accurate. The variable that contributes the most to our estimate of cyclists is date. Without the date predictor the correlation is reduced to half compared to the other experiments. However, when we eliminate the temperature predictor the correlation increases. Analysis. We have found a correlation between dates and bicycle flows. In addition we have been able to estimate the cycle flows, depending on date and weather. We did not expect that the variable temperature makes it harder for algorithms to estimate the number of cyclists. We assume that this is because people choose to cycle by date instead of the temperature.

Place, publisher, year, edition, pages
Malmö högskola/Teknik och samhälle , 2016. , p. 27
Keywords [sv]
Data Mining, Algorithm comparison, Estimation
National Category
Engineering and Technology
Identifiers
URN: urn:nbn:se:mau:diva-20129Local ID: 21194OAI: oai:DiVA.org:mau-20129DiVA, id: diva2:1479997
Educational program
TS Datavetenskap och applikationsutveckling
Available from: 2020-10-27 Created: 2020-10-27 Last updated: 2022-06-27Bibliographically approved

Open Access in DiVA

fulltext(364 kB)171 downloads
File information
File name FULLTEXT01.pdfFile size 364 kBChecksum SHA-512
d878abd390f3f99d72759e06442c046e7dfec13160c18fbae5ece26cd2ca709f1015b19ab1fd58a2d305eb9726154ebebe936bced0330c26250d4958810324e1
Type fulltextMimetype application/pdf

By organisation
Faculty of Technology and Society (TS)
Engineering and Technology

Search outside of DiVA

GoogleGoogle Scholar
Total: 171 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 312 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf