13.06.2023 aktualisiert

**** ******** ****
100 % verfügbar

Data scientist, data engineer, data analyst

Berlin, Deutschland
Nur Remote
Master in Information system and web
Berlin, Deutschland
Nur Remote
Master in Information system and web

Profilanlagen

CV - Henda Jami

Skills

Asp.NetHTMLJavaJavaScriptAPIsKünstliche IntelligenzApache AirflowData AnalysisKünstliche Neurale NetzwerkeAtlassian ConfluenceAtlassian JiraBusiness Process Model And NotationBusiness Process ManagementBusiness SoftwareC#C++CSSKrankenhausinformationssystemeClusteranalyseInformationssystemeComputerprogrammierungDatenbankenInformation EngineeringDateiWeb ScrapingETLDevOpsEclipseElasticsearchJ2EEEreignisprotokollierungGitHubHardware VirtualizationHibernateIntellij IDEAJava Database ConnectivityJava Persistence APIServletJSONWildflyjQueryPythonPostgreSQLLineare RegressionLogistische RegressionMachine LearningMicrosoft Sql-ServerMongoDBMonte-Carlo-SimulationMySQLNeo4jNltkNumPyOpen SourceWindows PowershellRecommendersystemeMicrosoft Power BITensorFlowSentiment DetectionServiceorientierte ArchitekturSQLSalesforce TableauUMLVersionierungVirtualBoxVirtualizationWeb ApplikationenAnacondaJupyter NotebookSupervised LearningChatbotsPyTorchFlaskRandom ForestVirtuelle UmgebungDeep LearningKaggleArchimateNaive BayesConvolutional Neural NetworksKerasGitpandasMatplotlibGitlab-CiPlotlyAtlassian BitbucketElastic KibanaDaten-PipelineDockerUnsupervised LearningProgramming Languages
ASP.NET, Airflow, Anaconda, API, ArchiMate, AI, Artificial Neural Network, Neural Network, Neural Networks, Bitbucket, Business Process Management, BPMN, Software Business, C#, C++, CSS, chatbot, clustering, programming, Atlassian Confluence, Confluence, Convolutional Neural Networks, data analytics, data analysis, data files, data pipeline, Databases, database, deep learning, DevOps, Docker, eclipse, ElasticSearch, event log, ETL, Flask, Git, GitHub, GitLab CI, HTML, Hardware Virtualization, Hibernate, Hospital Information Systems, user interface, web interface, data engineering, information system, Computer Science, IntelliJ Idea, Atlassian JIRA, JIRA, jQuery, JSON, JAVA, Java data base connectivity, JDBC, Java persistence API, JEE, Java Servlet, JAVA SCRIPT, Jupyter Notebook, Kaggle, Keras, Kibana, Linear regression, Logistic Regression, machine learning, Versioning, matplotlib, MS Access, Microsoft SQL Server, MS SQL Server, Microsoft SQL server 2005, word, MongoDB, Mongo DB, Monte-Carlo simulation, MySQL, MY SQL, NLTK, naïve bayes, Naive Bayes, Neo4J, numpy, open source, Opensource, Pandas, plotly, PostgreSQL, Postgres, PowerBI, Programming languages, Python, Pytorch, random forest, Recommender, SQL, SQL 2008, sentiment analysis, Service Oriented Architecture, Supervised Learning, Tableau, Tensorflow, UML, Unsupervised Learning, virtual environment, Oracle Virtualbox, Virtualization, web application, web app, web scraping, JBOSS, PowerShell

Sprachen

ArabischMutterspracheDeutschgutEnglischverhandlungssicherFranzösischMuttersprache

Projekthistorie

Software engineer

Higher Institute of Computer Science Kef Tunisia
Project Nr. 2 Service Oriented Architecture University project
Time Period 2018
Project duration 1 Months
Company/Institute Higher Institute of Computer Science Kef Tunisia
Sector Bank
Project team members 1
Project description Creation and deployment of a Service Oriented Architecture web
application
Position in Project Software engineer
Self implemented project parts The goal of this project was to design and develop an SOA JEE web
application used in Bank through the following steps:
- I developed a presentation tier (user interface) using HTML,
CSS
- I developed a web tier using Java Servlet
- I developed a business tier using Java persistence API
(Hibernate), Java data base connectivity JDBC
- I developed the enterprise information system (EIS) tier
using MY SQL
Hardware PC.
Software JBOSS, JEE, Java, My SQL, HTML, CSS, JDBC

Project Scientist

Higher Institute of Computer Science Kef Tunisia
Project Nr. 3 Multi Agent System University project
Time Period 2018
Project duration 1 Months
Company/Institute Higher Institute of Computer Science Kef Tunisia
Sector Multiagent System
Project team members 1
Project description Creation and deployment of a multi agent system using JADE
Position in Project Scientist
Self implemented project parts The goal of this project was to create and develop multi agents'
system using Java Agent Development Framework (JADE)
Hardware PC.
Software JADE, java

Data scientist

Hardware PC; Spiced Academy
Hardware PC
Software WSL, Anaconda, Python, Pandas, matplotlib, seaborn, sklearn,
pmdarima, GitHub




01-2023 Page 11




Project No. 13 Bootcamp Project-07
Time Period 09/2022 - 12/2022
Project duration 1 week
Company/Institute Spiced Academy
Sector Data Science
Project team members 1
Project description Predict customer behaviour in a supermarket, applying Markov
Chain modeling and Monte-Carlo simulation
Position in Project Data scientist
Self implemented project The goal of this project was to predict customer behavior in a
parts supermarket, applying Markov Chain modeling and Monte-Carlo
simulation.

The project included the following tasks:

Data Analysis Calculating Transition Probabilities between the aisles
Implementing a Customer Class Running MCMC (Markov-Chain
Monte-Carlo) simulation for a single class customer, through the
following steps:
- I explored the data (includes pandas wrangling)
- I calculated transition probabilities (a matrix) using
pd.crosstab()
- I implemented a customer class to run a MCMC simulation
for a single customer random .choices(states,
weights =probs[])
- I extended the simulation to multiple customers

Data Scientist

Hardware PC; Spiced Academy
Hardware PC
Software WSL, Anaconda, Python, Pandas, matplotlib, seaborn, sklearn,
random




01-2023 Page 12




Project No. 14 Bootcamp Project-08
Time Period 09/2022 - 12/2022
Project duration 1 week
Company/Institute Spiced Academy
Sector Data Science
Project team members 1
Project description Build an Artificial Neural Network that recognizes objects on images
made by the webcam
Position in Project Data Scientist
Self implemented project The goal of this project was to build an Artificial Neural Network that
parts recognizes objects on images made by the webcam:

The project included the following tasks:

Implementing a Feed-Forward Neural Network Backpropagation from
Scratch Building Neural Network with Keras Training Strategies /
Hyperparameters of Neural Networks Convolutional Neural Networks
(CNN) Classifying images made with webcam with Pre-trained
Networks (ResNet50, vgg16, mobilenet_v2) Image Detection,
through the following steps:

- I created a virtual environment for my project
- I collected an image dataset with my webcam
- I build a neural network from scratch keras .Sequential()
- I processed the dataset using ImageDataGenerator()
- I trained a neural network using a pre-built image dataset
- I predict a new image from my webcam using my model
- I used a pre-trained neural network to make predictions of
images from my webcam

Hardware PC
Software WSL, Anaconda, Python, Pandas, matplotlib, seaborn, sklearn,
tenserflow, keras

Data scientist

Spiced Academy
Project No. 07 Bootcamp Project-01
Time Period 09/2022 - 12/2022
Project duration 1 week
Company/Institute Spiced Academy
Sector Data Science
Project team members 1
Project description Visual data analysis animated scatterplot
Position in Project Data scientist
Self implemented project The goal of this project was to create an animated scatterplot (gif
parts file) to visualize the changes of countries' fertility rate, life expectancy
and population between 1960 and 2015 from a gapminder csv file
through the following steps:

Get data
- I downloaded gapminder csv data files
- I loaded data into pandas
- I inspected columns
Inspect variables
- I summarized / aggregated using descriptive statistics
- I drew some exploratory plots
Data wranglings
- I cleaned data and merged with other tables
- I rearranged axes
Visualize findings
- I Identified relevant facts
- I created explanatory plots: draw a histogram / scatterplot
using matplotlib and seaborn
- I presented key findings

Hardware PC
Software WSL, Anaconda, Python, Pandas, matplotlib, seaborn, GitHub

Data scientist

Spiced Academy
Project No. 08 Bootcamp Project-02
Time Period 09/2022 - 12/2022
Project duration 1 week
Company/Institute Spiced Academy
Sector Data Science
Project team members 1
Project description Classify the survival of Titanic
Position in Project Data scientist
Self implemented project The goal for this project was to build a machine learning model to
parts predict the survival of Titanic passenger based on the features in the
dataset of Kaggle's "Titanic - Machine Learning from Disaster"
through the following steps:

- I explored the Titanic dataset
- I split the data into a training and validation set using
train_test_split()
- I trained a Logistic Regression classification model using
LogisticRegression()
- I calculated the train and validation accuracy using
round(accuracy_score() )
- I processed data to enhance the score of the model by
applying some future engineering methods like
make_pipeline(),
SimpleImputer(),KBinsDiscretizer(),OneHotEncoder(),Standa
rdScaler(),ColumnTransformer()
- I trained a Random Forest classification model using
RandomForestClassifier()

Hardware PC
Software WSL, Anaconda, Python, Pandas, matplotlib, seaborn, sklearn

Data scientist

Spiced Academy
Project No. 09 Bootcamp Project-03
Time Period 09/2022 - 12/2022
Project duration 1 week
Company/Institute Spiced Academy
Sector Data Science
Project team members 1
Project description Regression bicycle rental prediction
Position in Project Data scientist
Self implemented project The goal for this project was to build and train a regression model on
parts the Capital Bike Share (Washington, D.C.) Kaggle data set, in order
to predict demand for bicycle rentals at any given hour, based on
time and weather, e.g.:", through the following steps:

- I explored the data set
- I split the data into a training and validation set
train_test_split(), and create time related features
- I trained a linear regression model using LinearRegression()
- I optimized the model by expanding or selecting features
- I regularized the model to avoid overfitting
- I calculated the Root Means Squared Log Error RMSLE for the
training and validation set using mean_squared_error()

Hardware PC
Software WSL, Anaconda, Python, Pandas, matplotlib, seaborn, sklearn

Data scientist

Spiced Academy
Project No. 10 Bootcamp Project-04
Time Period 09/2022 - 12/2022
Project duration 1 week
Company/Institute Spiced Academy
Sector Data Science
Project team members 1
Project description Build a text classification model on song lyrics
Position in Project Data scientist
Self implemented project The goal for this project was to build a text classification model on
parts song lyrics to predict the artist from a piece of text through the
following steps:

- I created my user data set (corpus) through web scraping
with BeautifulSoup, the song-lyrics of selected artists are
extracted from lyrics.com.
- I built a model pipeline using Tfidfvectorizer (TF-IDF) that
can transforms the words of the corpus into a matrix, countvectorizes
and normalizes them at once by default.
For classification, the multinomial Naive Bayes classifier
MultinomialNB() was used which is suitable for classification
with discrete features like word counts for text classification.

Hardware PC
Software WSL, Anaconda, Python, Pandas, matplotlib, seaborn, sklearn, bs4,
GitHub

Data scientist

Spiced Academy
Project No. 11 Bootcamp Project-05
Time Period 09/2022 - 12/2022
Project duration 1week
Company/Institute Spiced Academy
Sector Data Science
Project team members 1
Project description Build a dockerized data pipeline to analyze the sentiment of tweets
Position in Project Data scientist
Self implemented project The challenge of this Data Engineering project was to build a
parts Dockerized Data Pipeline to analyze the sentiment of tweets. At first,
using Tweepy API, tweets are collected in a selected topic and stored
in a MondoDB database (tweet_collector). Next, the sentiment of
tweets is analyzed and the tweets with the scores are stored in a
Postgres database (ETL_job). Finally, tweets with sentiment score are
published on a Slack channel. (slack_bot) through the following
steps:

- I installed Docker
- I build a data pipeline with docker-compose
- I collected Tweets from Tweepy API and store Tweets in
Mongo DB
- I created an ETL job transporting data from MongoDB to
PostgreSQL
- For the sentiment analysis, SentimentIntensityAnalyzer() of
the the Vader library (Valence Aware Dictionary and
sEntiment Reasoner) was used
- Build a Slack bot that publishes selected tweets using SQL


Hardware PC
Software WSL, Anaconda, Python, Pandas, matplotlib, seaborn, sklearn,
Docker,Docker compose, ETL, Tweepy API, MondoDB, PostgreSQL,
slack, SQL

Data scientist

Spiced Academy
Project No. 12 Bootcamp Project-06
Time Period 09/2022 - 12/2022
Project duration 1 week
Company/Institute Spiced Academy
Sector Data Science
Project team members 1
Project description Time series temperature forecast
Position in Project Data scientist
Self implemented project In this project, I applied the AutoRegressive Integrated Moving
parts Average ARIMA model for a short-term temperature forecast. After
visualizing the trend, the seasonality and the remainder of the time
series data (daily mean temperature in Berlin-Treptow from 1979-
2020), I run tests such as ADF and KPSS for checking stationarity
(time dependence).

For determining the parameters of the ARIMA model (p, d, q), I
present two approaches:

Inspecting the lags of the Autocorrelation Fuction (ACF) and Partial
Auto Correlation Functions (PACF) plots. Using alkaline-ml Auto-
Arima process which automatically finds the most optimal order
setting that has the lowest AIK, through the following steps:

- I got and cleaned temperature data from www.ecad.eu
- I build a baseline model modelling trend and seasonality
- I plotted and inspected the different components of a time
series
- I built a model time dependence of the remainder using an
AR model using pmdarima.AutoARIMA()
- Compare the statistical output of different AR models using
plot_acf, plot_pacf
- I predicted temperature for 5 next years

Data Scientist

Spiced Academy
Project No. 15 Bootcamp Project-09
Time Period 09/2022 - 12/2022
Project duration 1 week
Company/Institute Spiced Academy
Sector Data Science
Project team members 1
Project description Recommender movie system
Position in Project Data Scientist
Self implemented project The goal of this project was to build a proof of concept: A web
parts application that showcases different movie recommendation
algorithms., through the following steps:

- I downloaded a small version of the MovieLens-dataset
- I implement a baseline recommender using
recommend_random
- I derive a user-item matrix
- I picked and implemented a Collaborative Filtering
recommender: (Collaborative Filtering with Matrix
Factorization, Neighborhood based Collaborative Filtering)
using recommend_with_NMF
- I wrote a flask web interface
- I connected my recommender-model to flask using Flask,
render_template, request

Hardware PC
Software WSL, Anaconda, Python, Pandas, matplotlib, seaborn, sklearn,
Flask, recommender

Data Scientist

Spiced Academy
Project No. 16 Bootcamp Project-10
Time Period 09/2022 - 12/2022
Project duration 1 week
Company/Institute Spiced Academy
Sector Data Science
Project team members 1
Project description Flask chatbot
Position in Project Data Scientist
Self implemented project The goal of this project was to build a proof of concept: A web
parts application that chat with a robot using deep learning, through the
following steps:

- I downloaded a json dataset
- I used WorldNetLemmatizer() to create my corpus and clean
the extracted texts. During the text pre-processing, wordtokenizer
and word-lemmatizer of Natural Language Toolkit
(NLTK) is used
- I classifyed my data into 0's and 1's because neural networks
work with numerical values
- I built and deployed a Sequential CNN model, that we'll train
on the dataset we prepared above using keras .Sequential()
- I wrote a flask web interface
- I connected our chatbot to flask using Flask,
render_template, request

Hardware PC
Software WSL, Anaconda, Python, Pandas, matplotlib, seaborn, sklearn, nltk,
Flask

Data scientist

InfoarchiteQ-Labs GMBH
Project No. 5 InfoarchiteQ-Labs GMBH Internship project
Time Period 01/2022 - 06/2022
Project duration 3 months
Company/Institute InfoarchiteQ-Labs GMBH
Sector Consulting
Project team members 3
Project description Data science experiment project: COVID analytics
Position in Project Data scientist
Self implemented project The goal of this project was to analyze an COVID dataset and
parts presents keys findegs using jupyter notebook/pandas through the
following steps:

Get data
- I downloaded the data set
- I loaded data into pandas
- I inspected the columns
Inspect variables
- I summarized / aggregated data using descriptive statistics
- I drew some exploratory plots
Data wranglings
- I cleaned data and merged with other tables
- I rearranged axes
Visualize findings
- I Identified relevant facts
- I created explanatory plots: draw a histogram / geopandas
using matplotlib and seaborn
- I presented key findings

Hardware Virtualization
Software JIRA API, Confluence API, ElasticSearch, Juypter Notebook, Slack,
Anaconda, Python, Pandas, matplotlib, seaborn, geopandas

Data scientist

InfoarchiteQ-Labs GMBH
Project duration 3 months
Company/Institute InfoarchiteQ-Labs GMBH
Sector Consulting
Project team members 3
Project description Analyze an opensource license reporting based on the ScanCode
frameworks
Position in Project Data scientist
Self implemented project The goal of the proof-of-concept was to analyse an automated
parts capture of the integrated Opensource components using the
Opensource frameworks ScanCode Toolkit jupyter notebook and
Kibana.

I got the output (json file) of a Gitlab CI/CD pipeline programming
that identifies and references the existing opensource libraries within
software repositories using the ScanCode Toolkit.

I explored, cleaned, aggregated data with descriptive statistics using
Pandas, /Juypter-Notesbook and made available for a Kibana
dashboard.

Created a Kibana dashboard by answering some business goals.

Hardware PC
Software ScanCode Toolkit, Gitlab CI/CD, Elastic-stac, pandas, Kibana,
Juypter Notebook, Slack

Project Scientist

Higher Institute of Computer Science Kef Tunisia; Conservatoire National des Arts et des Metier
Project Nr. 4 Master thesis
Time Period 01/2019 - 06/2019
Project duration 6 Months
Company/Institute Higher Institute of Computer Science Kef Tunisia
Research Internship in Conservatoire National des Arts et des Metier
CNAM Paris France
Sector Healthcare analytics
Project team members 1
Project description Analyse of healthcare process using process mining technique.
Position in Project Scientist
Self implemented project parts The goal of my master thesis was to use process mining techniques
to analyse and extract knowledge from an event log file (Healthcare
process) produced by Hospital Information Systems where it is vital
that people can deviate to deal with changing circumstances.

I analysed a real-life Hospital log with same robust algorithms
existing in Prom framework.

I began by demonstrating Business Process Management BPM,
Process Mining PM and healthcare process, then I explored the event
logs using dottet chart technique.

I discovered healthcare model using Heuristics miner and I
discovered deviation cases and process variants using trace
clustering technique.

Later I implemented a new plug-in using Prom Toolkit open source
and eclipse (Java) to add new classifiers to the event log file based
on its attributes (timestamp) to improve and refine our analysis.
Finely I reanalysed the file transformed with the same process
mining techniques, and I compared results.

Hardware PC.
Software Business Process Management (BPM), Process Mining, dottet chart,
Heuristics miner, trace clustering, Prom Toolkit, eclipse, java

Software engineer

Higher Institute of Computer Science Kef Tunisia
Project Nr. 1 Bachelor project
Time Period 01/2009 - 06/2009
Project duration 6 Months
Company/Institute Higher Institute of Computer Science Kef Tunisia
Sector Industry
Project team members 1
Project description Creation and deployment of a web application used in a factory for
the traceability of product errors
Position in Project Software engineer
Self implemented project parts The goal of this project was to design and develop a three-tier web
application used in a factory for the traceability of product errors
through the following steps:
- I designed the web app using Open ModelSpher and UML
(use case diagram, sequence diagram, class diagram)
- I developed a presentation tier (user interface) using
ASP.NET
- I developed a business tier using c#
- I developed a data tier using Microsoft SQL server 2005
Hardware PC.
Software UML, Open ModelSphere, C#, Microsoft SQL server 2005, ASP.NET

Kontaktanfrage

Einloggen & anfragen.

Das Kontaktformular ist nur für eingeloggte Nutzer verfügbar.

RegistrierenAnmelden