1

Python Data Wrangler Internship Jobs in New York

... Python, SQL) - Machine learning & statistics - Data visualization (Tableau, Power BI, Matplotlib, Looker) - Data wrangling and preprocessing - Critical thinking and problem solving - Business acumen ...

... code for data wrangling, modeling, and AI-assisted analytical workflows. • Perform deep ... in Python, including analytical and modeling libraries. • Experience applying AI, machine ...

... code for data wrangling, modeling, and AI-assisted analytical workflows. • Perform deep ... in Python, including analytical and modeling libraries. • Experience applying AI, machine ...

Data Analyst

Iselin, NJ · On-site +1

$95K - $110K/yr

Strong SQL and Python skills; experience with data wrangling, transformation, and automation. * Familiarity with financial concepts such as revenue recognition, fee structures, and advisor ...

Senior Data Scientist

Newark, NJ · On-site

$97.30 - $178.80/hr

Data Wrangling: Preparing data for further analysis; Redefining and mapping raw data to generate ... Python, SQL What We Offer You Prudential is required by state specific laws to include the salary ...

Data Scientist

Manhattan, NY · On-site

$96K - $121K/yr

... like Python or R * 4+ years of experience with complex SQL * 3+ years of experience with the ... Experience with data wrangling techniques to cleanse data for data science applications

Your work will involve designing AI systems, data wrangling, and software implementation to enable ... such as Python and C++ to develop and deploy AI models - Managing complex data analysis and ...

Your work will involve designing AI systems, data wrangling, and software implementation to enable ... such as Python and C++ to develop and deploy AI models - Managing complex data analysis and ...

Your work will involve designing AI systems, data wrangling, and software implementation to enable ... such as Python and C++ to develop and deploy AI models - Managing complex data analysis and ...

Showing results 21-40

Python Data Wrangler Internship information

What is the difference between Python Data Wrangler Internship vs Data Analyst Intern?

AspectPython Data Wrangler InternshipData Analyst Intern
Required SkillsPython, data cleaning, scriptingExcel, SQL, basic statistics
Work EnvironmentData preprocessing, scripting tasksData analysis, reporting
Industry UsageTech, finance, startupsBusiness, marketing, finance

The Python Data Wrangler Internship focuses on data cleaning and scripting using Python, ideal for those interested in data preprocessing roles. In contrast, Data Analyst Internships involve analyzing data and creating reports, often requiring Excel and SQL skills. Both roles are common in tech and business sectors, but they serve different stages of data handling and analysis.

What are the key skills and qualifications needed to thrive as a Python Data Wrangler intern, and why are they important?

To thrive as a Python Data Wrangler Intern, you need a solid understanding of Python programming, data manipulation, and foundational knowledge in statistics or data science. Familiarity with tools such as pandas, NumPy, Jupyter Notebooks, and version control systems like Git is typically expected. Strong problem-solving, attention to detail, and effective communication skills set candidates apart in this role. These abilities are crucial for efficiently cleaning, transforming, and interpreting data to support accurate analysis and collaborative project work.

What are some typical projects or tasks assigned to Python Data Wrangler interns, and how do these contribute to the overall data team?

As a Python Data Wrangler intern, you are likely to work on projects involving data cleaning, transformation, and integration from multiple sources. Typical tasks include writing scripts to automate data preprocessing, identifying and correcting inconsistencies, and ensuring data integrity for analytics pipelines. You'll often collaborate closely with data analysts, data engineers, and sometimes machine learning teams to prepare datasets that drive business insights. These contributions are vital, as high-quality, well-structured data forms the foundation for all downstream analytics and decision-making within the organization.

What does a Python Data Wrangler intern do?

A Python Data Wrangler Intern is responsible for collecting, cleaning, organizing, and preparing data for analysis, primarily using Python programming. Their work involves writing scripts to handle messy or unstructured data from various sources, ensuring it is accurate and usable for data science or analytics teams. Tasks may include automating data extraction, transforming data formats, and collaborating with team members to support data-driven projects. This role provides hands-on experience with real-world datasets and tools commonly used in the industry.
What are popular job titles related to Python Data Wrangler Internship jobs in New York? For Python Data Wrangler Internship jobs in New York, the most frequently searched job titles are:
What job categories do people searching Python Data Wrangler Internship jobs in New York look for? The top searched job categories for Python Data Wrangler Internship jobs in New York are:
What cities in New York are hiring for Python Data Wrangler Internship jobs? Cities in New York with the most Python Data Wrangler Internship job openings:
Infographic showing various Python Data Wrangler Internship job openings in New York as of August 2026, with employment types broken down into 1% As Needed, 81% Full Time, 14% Part Time, 1% Temporary, and 3% Contract. Highlights an 86% Physical, 4% Hybrid, and 10% Remote job distribution.

Data Scientist, Portfolio Optimization

Formation Bio

New York, NY • On-site

Full-time

Re-posted 16 days ago


Job description

About Formation Bio
Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development.
Advancements in AI and drug discovery are creating more candidate drugs than the industry can progress because of the high cost and time of clinical trials. Recognizing that this development bottleneck may ultimately limit the number of new medicines that can reach patients, Formation Bio, founded in 2016 as TrialSpark Inc., has built technology platforms, processes, and capabilities to accelerate all aspects of drug development and clinical trials. Formation Bio partners, acquires, or in-licenses drugs from pharma companies, research organizations, and biotechs to develop programs past clinical proof of concept and beyond, ultimately helping to bring new medicines to patients. The company is backed by investors across pharma and tech, including a16z, Sequoia, Sanofi, Thrive Capital, John Doerr, Spark Capital, SV Angel Growth, and others.
You can read more at the following links:
  • Our Vision for AI in Pharma
  • Our Current Drug Portfolio
  • Our Technology & Platform

At Formation Bio, our values are the driving force behind our mission to revolutionize the pharma industry. Every team and individual at the company shares these same values, and every team and individual plays a key part in our mission to bring new treatments to patients faster and more efficiently.
About the Position
As a Data Scientist on the platform prediction team, you'll translate our probability of success predictions into measurable portfolio-level outcomes. You'll architect core systems - order management, execution simulation, portfolio construction, risk monitoring, and performance attribution - that let us rigorously evaluate signals from our AI-driven predictions in public and private equities and our internal portfolio.
This role sits at the intersection of quantitative finance, healthcare data, and AI-driven drug development. If you're excited about applying portfolio construction and risk management fundamentals to one of the most consequential prediction problems in healthcare, this is the role.No other company - hedge fund or pharma - has a technical data science position translating drug development experience into durable AI-native portfolio strategies. The skills you develop here - portfolio construction over assets with radically asymmetric risk profiles, clinical trial analytics, AI/ML in production, and risk management across multi-year horizons - can directly impact the delivery of new and effective therapeutics to patients by best aligning impactful medicines with economic incentives.
Responsibilities
  • Work with the team to implement and maintain core portfolio engine: order management system, execution simulation layer, portfolio construction service, and performance tracking
  • Design risk frameworks that quantify exposure across a portfolio of drug development bets with radically different risk profiles, timelines, and failure modes
  • Run rigorous backtesting experiments with strict temporal constraints to evaluate Formation strategies against baseline approaches and measure marginal signal from new evidence sources
  • Coordinate across the organization to integrate internal Formation data sources (clinical trial data, genomic evidence, real-world data) and proprietary tooling into portfolio analytics pipelines
  • Work with product and engineering teams to build dashboards and reporting that communicate portfolio performance, risk metrics, and strategy comparisons to both technical and executive stakeholders
  • Collaborate with the broader data science team to ensure portfolio-level evaluation feeds back into model improvement and evidence prioritization

About You
Required Qualifications
  • PhD in a quantitative field (statistics, finance, physics, computational science, engineering, or related)
  • 1-3 years in a quantitative research, data science, or analytics role in life sciences or life science adjacent field (healthcare, academic research, or consulting all count; substantive internships qualify)
  • Strong Python programming skills with experience in data-intensive workflows (pandas, numpy, scipy)
  • Solid grasp of core portfolio construction and risk concepts: position sizing, rebalancing, Sharpe ratio, drawdown, volatility, benchmark comparison
  • Demonstrated ability to work with messy, real-world datasets - comfortable with data wrangling, deduplication, and quality assessment
  • Clear communicator who can present quantitative results to both technical peers and business stakeholders

Preferred Qualifications
  • Experience with backtesting frameworks or portfolio simulation (vectorbt, Backtrader, or custom implementations)
  • Exposure to healthcare, pharma, or biotech data (clinical trials, claims data, -omics, real-world evidence)
  • Familiarity with alternative data in a research or investment context
  • Experience with probability-of-success modeling, drug development decision analysis, or health economics
  • Comfort with LLMs or AI/ML pipelines in a production or research setting
  • Familiarity with dashboard/visualization tools (Streamlit, Plotly, Dash) and pipeline orchestration (Dagster, Airflow)

Healthcare OR finance domain knowledge is valued; both are not required.
Total Compensation Range: $154,500 - $202,000
Compensation Individual compensation is determined by several factors, including role scope, geographic location, and skills & experience. Your offer will reflect where you fall within the range based on these considerations. In addition to base salary, we offer equity, comprehensive benefits, and generous perks. If the posted range doesn't match your expectations, we still encourage you to apply!
Where We Hire Formation Bio is prioritizing hiring in key hubs, primarily the New York City and Boston metro areas, with a hybrid model requiring 3 days per week in office. Applicants from the Research Triangle (NC) and San Francisco Bay Area may also be considered. Please apply only if you reside in these locations or are willing to relocate
Equal Opportunity Formation Bio is committed to building a diverse and inclusive team. We are an equal opportunity employer and welcome candidates from all backgrounds. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, national origin, ancestry, sex (including pregnancy, childbirth, breastfeeding, and related medical conditions), gender identity or expression, sexual orientation, age, disability, genetic information, marital status, military or veteran status, or any other characteristic protected by federal, state, or local law.